That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.
These are essentially zero-shot classifiers; they don't need to be trained for a specific classification task. You could include some natural language context on the rules for classification and it should get good enough accuracy.
This isn't really anything new it just seems like a new API but you could do the exact same thing with just a little bit of prompt engineering all the way back when GPT-3 was first released. Am I missing something?
No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.
For most use cases is that actually needed though? Just having it choose between predefined responses seems like enough but I'm curious about specific use cases because I do feel like I'm missing something
This is useful for classification problems; any time you need to write software that looks at some fuzzy data and needs to make a probabilistic decision. It's far more cost-efficient and performant to use this type of model instead of an LLM.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.
The response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.
Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.
If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
At my job people discuss almost weekly what the best models are for various price/quality points, with our CTO in particular wanting us to make sure we're being cost-effective in terms of what we're using for a given task (his words are basically "I don't want you to use these tools less, I just want you to use them effectively")
A lot of places already subscribe to both. Indeed, in large enterprises it isn’t uncommon to simultaneously subscribe to Copilot, Claude, OpenAI, Cursor, AWS Bedrock, Gemini, etc — you might subscribe to different ones for different teams/projects/employees/etc, but often the approval by legal/IT/etc is generic not scoped to whoever is using it right now
If you are charged based on usage, you can “soft switch” between them really quickly.
It's a problem for the security, legal, and IT teams, but not that much for developers, unless they get really particular about their harness. On the other hand, these are the high dollar value accounts that providers want to keep; but if there's no reason to avoid switching, this market really will feel like a utility market (i.e. it'll be like switching ISPs or cell phone providers - annoying but fungible).
The AI companies want to differentiate and become something more than a commodity, even if it's as critical as a utility is.
Enterprises tend to be your largest customers. Especially when current stock values for tech companies are majority predicated on AI becoming a staple in everyday life.
Agree. Switching models with a keypress is for developer coding.
Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.
For every production use case I build a eval suite which I use for prompt tuning and model evaluation and config. How else do you establish your model and prompt combination works? How do you upgrade to a newer model or decide on a fallback?
This same suite can simply be run very handoff to switch to a new prod model.
My gut is that the market for "ai as a tool" aka Claude Code/Codex/computer use/etc is a significantly bigger one than "models behind the scenes of some service". I've seen people use evals a lot in the latter case and very little in the former case (outside of people's whose job is basically to review the new releases). I haven't personally met anyone with something like "here are a bunch of tickets + a snapshot repo checkout, please try to solve them all" eval approach.
Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"
Not only can you switch models easily, but their competition can easily duplicate their product. I'm not sure stickiness has been found yet, but my guess is being more of G-Suite for "intelligence" than being a token hawker.
If you are a business dealing with with anything remotely sensitive then this is not so easy and you are basically forced to do business with a big player.
> If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
Hopefully people will flock to whatever product is making its mission to be commodity and the easiest to replace. Really don't want another free ingress, 100$/TB egress Cloud situation.
But OpenAI and Ant can be in more than just the commodity part of the business. They can both make frontier models and sticky products on top of those that have network or other effects that make competition tough.
The question isn't whether the market is big enough. It's whether at a race to the bottom on prices is sustainable and worth what investors expected it to be. The original story of OpenAI and Anthropic was that it was a race to super intelligence and whoever gets there first basically captures all the money in the economy. That kind of investment pitch means missing out not just on big returns but perhaps infinite return on investment. It seemed like it'd be true when the market was young and the cost to compete was prohibitive, but now anyone can compete with open weight models that are cheaper and maybe not the best, but are good enough.
You raise $200B to be a high margin low capex business, not an industrial commodity producer with high capex and margins being set by competitors who can duplicate your product and undercut you on price.
Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).
Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).
Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).
Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.
Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).
Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.
Well, Jev doesn't meet any real compliance requirements but OpenAI's models do. So even if Jev is faster, any customers with compliance needs will obviously pick OpenAI because they can't pick Jev out of necessity.
I think releasing something like this makes sense even if it's underbaked, it's still very cheap, and if you have existing enterprise OpenAI relationship it's a lot easier to onboard something like this than set up a new Jev contract.
How these models play out is an open question but existing provider contracts and T&C are important for enterprise.
Good point, commercially, being an existing partner is always easier for adoption. Heck, why I'd like to get Mercury Decide to replace Jev myself, rather than one than two to work with.
Still surprised they even leveraged Luna for this. Given their resources in data, compute and manpower, would training a decision model from scratch take that much longer to not make sense given the cost, compute and performance advantages that would likely provide?
3 times more expensive at twice the latency with lower performance is a tough sell, though yeah, prior relationships will likely smooth some of those deficiencies over.
yeah and this isn't a long term solution, stand this up, see what value you get out of it, and in a few months you cans witch to whatever the best decision model is
They can follow up in N weeks with a better one. Even if your eval is true and it’s worse, planting a flag makes sense. Some people will just use OAI because it’s OAI. No one will remember their week 2 evals in a few months.
i would be shocked if luna decides is less generally capable than jev. jev has failed to understand any novel domain I've given it. i have found that i use it only when "some data is better than no data"
Very task-dependent of course and mine are unique to say the least, so could see Luna being better in certain domains, even if my measurements have not shown that yet, happy for anyone to show otherwise.
For what it's worth, ran every task twice on each model, most were for some UI component synthesis and charting insanity that is a bit hard to explain, but some were simple tag selection, basic noul at threshold 60%. Essentially, whether to use the provided tag given the title of a browser tile:
Of course, tags can be a bit subjective, but in these cases, I'd argue the values provided by Jev were far more representative of my subjective assessment over Lunas. If SnP stuff on Bloomberg isn't investing, nothing is.
Goal for tagging is mainly a near instant, over writable, sane default provided to users in the background. Resolve the whole "I love using Notion/Obsidian/PKM software of your choice but spend 80% of my time just thinking about the ideal tag before starting to read" issue. Lunas output is not really helpful here.
1. Title: "Mortgage calculator: estimate your monthly payment (Bankrate)"
Tag: "house hunting"
Decisions API (2 runs): 0.99, 0.99
2. Title: "S&P 500 index: live chart and news (Bloomberg)"
Tag: "investing"
Decisions API (2 runs): 1.0, 1.0
If you have other examples of requests with unexpected outputs, feel free to email me at by@openai.com and we can try to get to the bottom of it. Thanks for trying out the API!
Note that, being that this is gpt-6-luna under the hood, this offers you 1m token input window, and multi-modal (image) input. In my testing so far, I'm seeing 160-175ms end to end. Worst 5% 285ms, worst so far was 743ms.
All this is for extraordinarily simple decisions. Real world problems often are a lot more complex requiring highly structured outputs covering many output attributes and substructures, for which a conventional structured output via a documented schema is better. If instead you make twenty independent calls to a decisions API, you lose coherence among your twenty decisions. I think any hype surrounding decisions will be forgotten soon enough.
Deterministic predictions are sought by those looking to offload their decision responsibility to AI, whether to lower perceived legal risk or otherwise. This is fake risk reduction, i.e. "risk theater".
Unfortunately, a deterministic prediction utterly fails to yield an uncertainty measurement which is critical to have in actual risk reduction. If you want the variance in measurement, it is vital to obtain multiple measurements. This also gives a confidence interval.
It seems this one caught openai on the back foot, and this is a scramble to maintain parity.
The company is clearly still innovating towards AGI rather than asking "what do people actually need?"
Despite once being the darling of AI it's
- lost it's models' performance edge, and got too many similar offerings
- continues to launch products without a market or isn't done better using other tools (e.g. dots)
- despite having AI can't lock down it's own products showing lack of skill
- focuses on solving maths problems humans can do for tests, when real world problems - disease, materials, energy research etc is all outstanding
- abandoned it's open model and open source programmes, despite Google, and multiple successful Chinese, and now European companies make their frontier models open weight.
- pissed off a portion of its non corporate fan base by killing GPT 4o instead of recognising the brand and product attachment as an opportunity
- launched laughing stock projects like being able to actually call chatgpt on a telephone number (wtf!)
- fails to capitalise on market segments like an AI that can provide corporate network sentry duties
It's increasingly looking like the company has jumped the shark and if I was an investor would be asking questions as to why it actually took so long to bring a jev like product to market, and why they are labelling something that is a simplification of existing models as "beta".
The whole point of AI as I see it is to make our life easier and answer the questions we can't. It isn't to make an AGI so powerful that it can replace us.
Along the way, that goal was forgotten, but it's not been forgotten by the new startups.
Have there been any signals from Anthropic about matching this? We use AWS bedrock and just switched to Anthropic from OpenAI because of the ZDR guarantee. Would be great to not have to entertain switching back.
Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API
One of the examples on the docs page is it playing a video game. Doubt it’ll be able to run anything complex though. You’re simply trading accuracy for speed.
Every time something comes during the AI bubble we get a wave of me-toos. I think the Jev wave is noticeable for how muted it is.
The fact that this keeps happening demonstrates there is no moat. The fact that each wave gets a little less attention demonstrates there is no killer product here.
This rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.
Interesting to see where all this leads us and if other major labs follow suit
Edit: decisions voice looks really interesting as well[1]
Decisions voice isn't a product for anyone else who was confused: it's a canned guide for hooking up a voice model to the decision model browser use thing
> The tulip became a luxury item and many varieties were introduced. The varieties were classified and the most sought-after, prized tulips were the streaked tulips, especially yellow or white streaks on a red or purple background. These flame-like tulips were highly sought after. Interestingly, the streaks or “flames” of the tulip petals were caused by a virus. The virus is the tulip breaking virus, or tulip mosaic virus.
I genuinely do not understand why anyone would pay OpenAI for this. Running something comparable to Jev is pretty trivial. The whole point of paying for ChatGPT is because OpenAI has a bunch of warehouses that can run a zillion-parameter model.
Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.
Yeah and OAI is twice as expensive as Jev, which is kind of my point. And more expensive than open models, which you don't necessarily have to host yourself. Pure bandwaggoning.
For my use case it will cost like $11 a month and we already have OpenaAI keys and accounts with billing in place. I don't want to run my own model infra and I don't want to get permission to set up an account with typesafe.ai
If you're in an enterprise that already has a procurement agreement with OpenAI, this means you don't have to onboard another vendor. Bucket platform strategy.
My opinion is similar, but for a different reason: every use case for decision models that I can think of, I don’t want the model to change in X weeks when the lab decides to “improve it” or “make it safer”.
Depends on the quality of the results. These things are driven by text prompts. If it turns out the OpenAI one returns better quality results than open weight variants they'll be rewarded by the market.
Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.
There isn’t a moat in the sense of self hosting but you need a reason for people who don’t want that to stay on your platform. Customers save time and effort managing payments easier this way. However it’s a race to the bottom price wise.
Going to be all about branding and platform stickiness for OpenAI to make investors and creditors whole.
Existing enterprise contracts? Data retention contracts (some have zero data retention contracts)? Staying with a single provider because it's easier to have everything in one place?
If you work for a company that has a 3 to 6 month onboarding period for new vendors and a lifetime commitment to maintain a whole bunch of vendor management horseshit for as long as that relationship exists, it makes a ton of sense.
Add in a bunch of model governance and oversight for anything you train yourself and it’s pretty much a slam dunk deal.
(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)
>defecto standard for other providers.
I know it's just a little typo but it made my morning :)
How is this different from the categorisation models from ML era?
These are essentially zero-shot classifiers; they don't need to be trained for a specific classification task. You could include some natural language context on the rules for classification and it should get good enough accuracy.
The pitch is that it's a fully-general model, so you can skip training/tuning/selecting a particular categorisation model for each task.
That's great! So someone finally built the zero-shot model from the sales decks of 2015 =D
Specifically, it happened a few weeks ago when Typesafe released Jev; this is OpenAI's competitor to Typesafe.
This isn't really anything new it just seems like a new API but you could do the exact same thing with just a little bit of prompt engineering all the way back when GPT-3 was first released. Am I missing something?
No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.
For most use cases is that actually needed though? Just having it choose between predefined responses seems like enough but I'm curious about specific use cases because I do feel like I'm missing something
This is useful for classification problems; any time you need to write software that looks at some fuzzy data and needs to make a probabilistic decision. It's far more cost-efficient and performant to use this type of model instead of an LLM.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.
It's much faster and cheaper (an order of magnitude).
And theoretically will give you better answers statistically as it's calibrated.
The response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.
Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.
If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
Is making their product sticky perhaps what all the talk of supply chain security is really motivated by?
It takes me all of 2 keypresses to switch models. I don't know of a less sticky product
You've not seen how long it takes to switch an enterprise claude subscription to github copilot or vice versa with all the compliance and shareholders
Not sure where you work but at ours we have api access to all of them and can freely switch.
At my job people discuss almost weekly what the best models are for various price/quality points, with our CTO in particular wanting us to make sure we're being cost-effective in terms of what we're using for a given task (his words are basically "I don't want you to use these tools less, I just want you to use them effectively")
No DLP concerns?
I think they have zero data retention contracts with all of them.
Yes, they'll throw it out after they trained on it ¯\_(ツ)_/¯
I’ve seen enterprises have access to all the harnesses but not all the models within them
and they ask dumb followup questions after 7 business days when you want different access
A lot of places already subscribe to both. Indeed, in large enterprises it isn’t uncommon to simultaneously subscribe to Copilot, Claude, OpenAI, Cursor, AWS Bedrock, Gemini, etc — you might subscribe to different ones for different teams/projects/employees/etc, but often the approval by legal/IT/etc is generic not scoped to whoever is using it right now
If you are charged based on usage, you can “soft switch” between them really quickly.
A lot of places just pay the token cost of the usage, so they sign up to all of them and lets the winner win.
Comes to something then OpenAI’s best hope is to essentially become the next Oracle.
Lol at my $job we have one token quota which can be used with all the frontier models.
isn't that a problem for large enterprises?
It's a problem for the security, legal, and IT teams, but not that much for developers, unless they get really particular about their harness. On the other hand, these are the high dollar value accounts that providers want to keep; but if there's no reason to avoid switching, this market really will feel like a utility market (i.e. it'll be like switching ISPs or cell phone providers - annoying but fungible).
The AI companies want to differentiate and become something more than a commodity, even if it's as critical as a utility is.
Enterprises tend to be your largest customers. Especially when current stock values for tech companies are majority predicated on AI becoming a staple in everyday life.
This implies you didn’t run any sort of evaluations? It is not realistic for any sort of production use case to do this.
Agree. Switching models with a keypress is for developer coding.
Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.
For every production use case I build a eval suite which I use for prompt tuning and model evaluation and config. How else do you establish your model and prompt combination works? How do you upgrade to a newer model or decide on a fallback?
This same suite can simply be run very handoff to switch to a new prod model.
My gut is that the market for "ai as a tool" aka Claude Code/Codex/computer use/etc is a significantly bigger one than "models behind the scenes of some service". I've seen people use evals a lot in the latter case and very little in the former case (outside of people's whose job is basically to review the new releases). I haven't personally met anyone with something like "here are a bunch of tickets + a snapshot repo checkout, please try to solve them all" eval approach.
Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"
Not only can you switch models easily, but their competition can easily duplicate their product. I'm not sure stickiness has been found yet, but my guess is being more of G-Suite for "intelligence" than being a token hawker.
If you are a business dealing with with anything remotely sensitive then this is not so easy and you are basically forced to do business with a big player.
> If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
Hopefully people will flock to whatever product is making its mission to be commodity and the easiest to replace. Really don't want another free ingress, 100$/TB egress Cloud situation.
But OpenAI and Ant can be in more than just the commodity part of the business. They can both make frontier models and sticky products on top of those that have network or other effects that make competition tough.
The question isn't whether the market is big enough. It's whether at a race to the bottom on prices is sustainable and worth what investors expected it to be. The original story of OpenAI and Anthropic was that it was a race to super intelligence and whoever gets there first basically captures all the money in the economy. That kind of investment pitch means missing out not just on big returns but perhaps infinite return on investment. It seemed like it'd be true when the market was young and the cost to compete was prohibitive, but now anyone can compete with open weight models that are cheaper and maybe not the best, but are good enough.
You raise $200B to be a high margin low capex business, not an industrial commodity producer with high capex and margins being set by competitors who can duplicate your product and undercut you on price.
They are still gunning for ASI.
Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).
Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).
Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).
Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.
Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).
Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.
Well, Jev doesn't meet any real compliance requirements but OpenAI's models do. So even if Jev is faster, any customers with compliance needs will obviously pick OpenAI because they can't pick Jev out of necessity.
Or even just- they've already procured OpenAI.
That's a big motivator.
I think releasing something like this makes sense even if it's underbaked, it's still very cheap, and if you have existing enterprise OpenAI relationship it's a lot easier to onboard something like this than set up a new Jev contract.
How these models play out is an open question but existing provider contracts and T&C are important for enterprise.
Good point, commercially, being an existing partner is always easier for adoption. Heck, why I'd like to get Mercury Decide to replace Jev myself, rather than one than two to work with.
Still surprised they even leveraged Luna for this. Given their resources in data, compute and manpower, would training a decision model from scratch take that much longer to not make sense given the cost, compute and performance advantages that would likely provide?
3 times more expensive at twice the latency with lower performance is a tough sell, though yeah, prior relationships will likely smooth some of those deficiencies over.
yeah and this isn't a long term solution, stand this up, see what value you get out of it, and in a few months you cans witch to whatever the best decision model is
They can follow up in N weeks with a better one. Even if your eval is true and it’s worse, planting a flag makes sense. Some people will just use OAI because it’s OAI. No one will remember their week 2 evals in a few months.
i would be shocked if luna decides is less generally capable than jev. jev has failed to understand any novel domain I've given it. i have found that i use it only when "some data is better than no data"
Very task-dependent of course and mine are unique to say the least, so could see Luna being better in certain domains, even if my measurements have not shown that yet, happy for anyone to show otherwise.
For what it's worth, ran every task twice on each model, most were for some UI component synthesis and charting insanity that is a bit hard to explain, but some were simple tag selection, basic noul at threshold 60%. Essentially, whether to use the provided tag given the title of a browser tile:
Of course, tags can be a bit subjective, but in these cases, I'd argue the values provided by Jev were far more representative of my subjective assessment over Lunas. If SnP stuff on Bloomberg isn't investing, nothing is.Goal for tagging is mainly a near instant, over writable, sane default provided to users in the background. Resolve the whole "I love using Notion/Obsidian/PKM software of your choice but spend 80% of my time just thinking about the ideal tag before starting to read" issue. Lunas output is not really helpful here.
I tried reproducing your examples as predicate questions on the Decisions API (https://gist.github.com/by-openai/7b376d7866e36a4f38581f73ce...) and got:
If you have other examples of requests with unexpected outputs, feel free to email me at by@openai.com and we can try to get to the bottom of it. Thanks for trying out the API!Boy am I glad we're already dropping the "noul" term for a binary decision
Just going to drop this here: https://jeffyclassify.com/
Open source classifier models you can run and train locally on CPU
Note that, being that this is gpt-6-luna under the hood, this offers you 1m token input window, and multi-modal (image) input. In my testing so far, I'm seeing 160-175ms end to end. Worst 5% 285ms, worst so far was 743ms.
Image input? Wait a second; so this can be used for image classification without any kind of training at those speeds?
Yes! Check out the 20 second mark of this video for a demo of how fast image tasks can be: https://x.com/OpenAIDevs/status/2107573382229188645
If we compare this with using the older solution of writing a prompt to find out the answer of the classification
- Cost : It is the same for both scenarios $0.10 per 1M tokens
- Speed : decisions is 10x faster than responses API
- Quality : I guess if we compare with luna which is a pretty good model it itself, both will be at par
So essentially it has to do more with speed vs any other factor.
With decisions models, you don't pay for output tokens. Also this OpenAI API is multimodal.
Output tokens anyway will be minimal in a decision scenario, so even if you use responses API the cost will be less.
Not exactly. System 2 style reasoning to follow instructions can be expensive and counts as output tokens.
Jev really shook up the industry. This seems obvious in hindsight
All this is for extraordinarily simple decisions. Real world problems often are a lot more complex requiring highly structured outputs covering many output attributes and substructures, for which a conventional structured output via a documented schema is better. If instead you make twenty independent calls to a decisions API, you lose coherence among your twenty decisions. I think any hype surrounding decisions will be forgotten soon enough.
I don’t think so. People want more determinism and this is just another step in that direction.
Can't wait to read about BERT on the front-page!
Deterministic predictions are sought by those looking to offload their decision responsibility to AI, whether to lower perceived legal risk or otherwise. This is fake risk reduction, i.e. "risk theater".
Unfortunately, a deterministic prediction utterly fails to yield an uncertainty measurement which is critical to have in actual risk reduction. If you want the variance in measurement, it is vital to obtain multiple measurements. This also gives a confidence interval.
It’s interesting they skipped caching. I could see wanting to ask follow up questions so having your first x tokens in cache would be interesting.
Also if you have a long “system prompt” then caching would have saved a considerable amount on bulk data processing.
There may well be a technical reason I don’t understand.
I'd imagine it's chasing the lowest possible latency
Wouldn't caching be in favor of lowest possible latency?
It might be faster to burn the compute instead of having to fetch the KV-cache over the network from an SSD.
If you rather run your decision model on your CPU, check gutsy [0]
[0] - https://news.ycombinator.com/item?id=49976996
To quote Bruce Willis "Welcome to the party pal"
It seems this one caught openai on the back foot, and this is a scramble to maintain parity.
The company is clearly still innovating towards AGI rather than asking "what do people actually need?"
Despite once being the darling of AI it's
- lost it's models' performance edge, and got too many similar offerings
- continues to launch products without a market or isn't done better using other tools (e.g. dots)
- despite having AI can't lock down it's own products showing lack of skill
- focuses on solving maths problems humans can do for tests, when real world problems - disease, materials, energy research etc is all outstanding
- abandoned it's open model and open source programmes, despite Google, and multiple successful Chinese, and now European companies make their frontier models open weight.
- pissed off a portion of its non corporate fan base by killing GPT 4o instead of recognising the brand and product attachment as an opportunity
- launched laughing stock projects like being able to actually call chatgpt on a telephone number (wtf!)
- fails to capitalise on market segments like an AI that can provide corporate network sentry duties
It's increasingly looking like the company has jumped the shark and if I was an investor would be asking questions as to why it actually took so long to bring a jev like product to market, and why they are labelling something that is a simplification of existing models as "beta".
The whole point of AI as I see it is to make our life easier and answer the questions we can't. It isn't to make an AGI so powerful that it can replace us.
Along the way, that goal was forgotten, but it's not been forgotten by the new startups.
Have there been any signals from Anthropic about matching this? We use AWS bedrock and just switched to Anthropic from OpenAI because of the ZDR guarantee. Would be great to not have to entertain switching back.
> just switched to Anthropic from OpenAI because of the ZDR guarantee
Did you get that reversed? OpenAI has a ZDR guarantee while Anthropic doesn't.
Feels antithetical to their big model bitter lesson strategy
Why not try Jev?
Is there a mention of context length? I can't find it. I could not integrate Jev due to limited window (64k?).
Someone said 1mm
Overpriced crap, local models are better than this, it's also dumber than Luna for some reason, and Jev is definitely ~2-3x cheaper than this.
I am sure people will be able to use it, but if LLM progress is anything like before we will have Jev 2 in about a month.
Training a local model like Jev with some learnings that can be extremely cheap to host shouldn't take that long either.
So I honestly don't see the point of this, other than to put something out.
Although that does seem like OpenAI's strength turn around slop products and iteratively improve and try to out compete others in everything.
Only 2 things they have clearly given up on are Video models(no moat, copyright nightmare) and Music.
And it makes sense why. I feel like they will compete with even the no-name Dog, if the Dog launched a successful marketing video of an AI product.
I have seen this often in SF startups, heck I work for them, but man this is extreme.
But honestly all I see are long term price wars, I don't understand how this is a sustainable business strategy.
Or maybe that's the point... Who knows.
it already supports image inputs, which was the first big gap I found in Jev.
have you tried Moondream?
https://docs.moondream.ai/
Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API
One of the examples on the docs page is it playing a video game. Doubt it’ll be able to run anything complex though. You’re simply trading accuracy for speed.
Get an LLM to build it and report back! Cool idea
One difference between Decisions and Jev (for now) seems to be that Decisions can take image inputs, which is a pretty common need.
Clef can also do this.
How well calibrated is it? Is 90% calibrated to be correct 9 times out of 10?
I think Jev had put significant effort here and its not clear if luna will be well calibrated in this way.
Different APIs for different things reminds me of the early auto-complete vs instruction apis.
Will this get folded into models / post training pipelines at some point and make them better at calibrated outputs?
Every time something comes during the AI bubble we get a wave of me-toos. I think the Jev wave is noticeable for how muted it is.
The fact that this keeps happening demonstrates there is no moat. The fact that each wave gets a little less attention demonstrates there is no killer product here.
I wonder why the decision routing isn't just integrated into all models in addition to this stand alone.
What value is there in knowing if a can has a dent ?
Manufacturing quality control. There are automated tools to pull things like dented cans off a line.
So AI is a trillion dollar industry and one of the biggest thing it can do is spot dents on a can ?
This rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.
Interesting to see where all this leads us and if other major labs follow suit
Edit: decisions voice looks really interesting as well[1]
[0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev
[1]: https://developers.openai.com/api/docs/guides/decisions-voic...
Decisions voice isn't a product for anyone else who was confused: it's a canned guide for hooking up a voice model to the decision model browser use thing
But... Can it draw a pelican on bicycles?
How is the pricing vs Jev?
$0.10/mm input vs. $0.042/mm input. Both free output.
In the same bench a full Jev run cost USD 0.0192,- vs Luna at USD 0.06,-, both via OpenRouter today. So about 3x in favour of Jev.
> The tulip became a luxury item and many varieties were introduced. The varieties were classified and the most sought-after, prized tulips were the streaked tulips, especially yellow or white streaks on a red or purple background. These flame-like tulips were highly sought after. Interestingly, the streaks or “flames” of the tulip petals were caused by a virus. The virus is the tulip breaking virus, or tulip mosaic virus.
Source: https://www.canr.msu.edu/news/tulip_mania_the_history_of_the...
What's old is new.
You knew it was going to happen! Benchmarks or it didn't happen.
v3.26.0 of the openai Python SDK covers its use. Those already using the SDK don't need to make explicit HTTP calls.
Can we use this through subscription?
No. The OpenAI subscription has never covered any API calls. The closest you can probably use via subscription is to get structured outputs via Codex.
I genuinely do not understand why anyone would pay OpenAI for this. Running something comparable to Jev is pretty trivial. The whole point of paying for ChatGPT is because OpenAI has a bunch of warehouses that can run a zillion-parameter model.
Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.
Why would I run it myself? It's $0.10 per million tokens. Dirt cheap. (Jev is even cheaper.)
You could ask the same question about why anyone would rent a VPS. I can just run my own hardware, it's just a computer!
Buy vs rent is not just about what's possible, it's about what's economic.
For one, to remove network RTT
Yeah and OAI is twice as expensive as Jev, which is kind of my point. And more expensive than open models, which you don't necessarily have to host yourself. Pure bandwaggoning.
For my use case it will cost like $11 a month and we already have OpenaAI keys and accounts with billing in place. I don't want to run my own model infra and I don't want to get permission to set up an account with typesafe.ai
If you're in an enterprise that already has a procurement agreement with OpenAI, this means you don't have to onboard another vendor. Bucket platform strategy.
My opinion is similar, but for a different reason: every use case for decision models that I can think of, I don’t want the model to change in X weeks when the lab decides to “improve it” or “make it safer”.
Depends on the quality of the results. These things are driven by text prompts. If it turns out the OpenAI one returns better quality results than open weight variants they'll be rewarded by the market.
Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.
There isn’t a moat in the sense of self hosting but you need a reason for people who don’t want that to stay on your platform. Customers save time and effort managing payments easier this way. However it’s a race to the bottom price wise.
Going to be all about branding and platform stickiness for OpenAI to make investors and creditors whole.
Existing enterprise contracts? Data retention contracts (some have zero data retention contracts)? Staying with a single provider because it's easier to have everything in one place?
There are probably a lot more reasons.
If you work for a company that has a 3 to 6 month onboarding period for new vendors and a lifetime commitment to maintain a whole bunch of vendor management horseshit for as long as that relationship exists, it makes a ton of sense.
Add in a bunch of model governance and oversight for anything you train yourself and it’s pretty much a slam dunk deal.