Decision model releases from AWS, Cloudflare and OpenAI arrived within days of each other at the start of October, and on 9 October InfoWorld reported that the pattern now looks like a new tier of the enterprise AI stack. The idea is simple. The bounded choices an agent makes all day go to a small model built only to choose. Which tool should it call? Which team gets this ticket? Does this action need a human? A large reasoning model keeps the planning, the coding and the writing.
That split promises smaller token bills and faster AI agents. The analysts InfoWorld spoke to also warn of a new kind of AI-stack sprawl. Enterprises get more models to evaluate, confidence scores that mean different things from vendor to vendor, and hundreds of decision schemas that encode business policy without a clear owner.
This article sets out what each vendor shipped, how a decision model differs from the large language model above it, and what the new layer costs and saves. It then covers the calibration and governance work that decides whether the savings survive contact with production. It follows our coverage of the Jev model launch in September, and focuses on what InfoWorld’s analysts say the new layer costs to run.
Table of contents
- What InfoWorld Reported About the Decision Model Layer
- How a Decision Model Differs From the LLM Above It
- Where a Decision Model Sits in an Agent Stack
- The Price and Latency Case for a Decision Model
- AI-Stack Sprawl: The Hidden Cost of a Decision Model Layer
- Why One Decision Model’s 0.8 Is Not Another’s
- Decision Schemas Are Business Policy
- Cost per Successful Decision, Not Cost per Token
- Build, Rent or Abstract the Decision Model Layer
- A 90-Day Plan for Adding a Decision Model Layer
- Questions Buyers Are Asking About the Decision Model Layer
- References
What InfoWorld Reported About the Decision Model Layer
InfoWorld’s Anirban Ghoshal framed the trend around cost. Enterprises are trying to balance AI budgets against the demands of scaling agentic applications. They have tried hard-coding deterministic decisions and using smaller models for specific tasks. TypeSafe’s Jev, released in mid-September, added a third option: a specialised decision model for the bounded decisions that sit between an agent’s reasoning and its actions.
In the last week of September and the first days of October, three large vendors followed. Cloudflare released Clef and Clef-flash. AWS released Strands Decider 2B. OpenAI released a Decisions API that hides the model behind an endpoint. InfoWorld’s conclusion was that decision-making “could become another specialized layer in the enterprise AI stack”.
Six decision model launches in under three weeks
The table below lists what shipped, using each vendor’s own documentation. TypeSafe and Fastino had already started the category by the time the hyperscalers arrived, and Databricks added a SQL function in the same week as OpenAI. Our earlier piece on the wave of Jev alternatives covers the smaller open-source copies in more detail.
| Vendor and model | Released | Size and licence | List price (input) | Inputs |
|---|---|---|---|---|
| TypeSafe Jev 1.13 | 15 Sep 2026 | Size not published; hosted API | $0.042 per million tokens; output free | Text only |
| Fastino GLiNER2.5-Decide | 24 Sep 2026 | 340M encoder; Apache 2.0 | Open weights; your hardware | Text |
| Databricks ai_decide | End of Sep 2026 (beta) | Apache 2.0 models chosen by Databricks | Databricks SQL pricing | Text or structured data |
| OpenAI Decisions API | 30 Sep 2026 (public beta) | gpt-6-luna; model hidden | $0.10 per million tokens; no output charge | Text and images |
| Cloudflare Clef / Clef-flash | 1 Oct 2026 | 27B / 9B; Apache 2.0 weights | $0.24 / $0.09 per million tokens | Text, JSON, images, video |
| AWS Strands Decider 2B | 1 Oct 2026 | 2B; open source, data and scripts included | Open weights; your hardware | Text |
Why InfoWorld calls it a layer
The word “layer” is doing real work in the headline. AWS describes Strands Decider as a control-flow component that picks tools, routes tasks and decides an agent’s next step, which InfoWorld summarised as “in effect acting as an orchestrator for what an agent does next”. Cloudflare’s launch post suggests putting Clef “into the hot path for agents to make decisions” and pairing it with a larger model on Workers AI to take the action.
So the decision model is not a cheaper chatbot. It sits between reasoning and action and turns a fuzzy situation into a typed answer that ordinary code can branch on. David Linthicum, quoted in InfoWorld’s earlier coverage of Jev, put the old approach bluntly: using a general-purpose LLM for every bounded choice “is like using a full enterprise service bus to answer a yes/no routing question”.
How a Decision Model Differs From the LLM Above It
A decision model does not write. It reads a state (a support ticket, an agent transcript, a product photo) plus a set of typed questions, and returns a probability for every allowed answer. Because there is no free text to parse, the output is always one of the options you defined.
One pass, no generated text
The vendors use different architectures to get the same effect. Cloudflare runs its Qwen backbone in a single prefill-only pass and then scores the valid schema choices in parallel, so “there’s no intermediate text to generate token by token”. AWS takes a Qwen3.5-2B model, removes the head that generates text and replaces it with a pointer head of just over a million parameters that scores each option. Fastino’s GLiNER2.5-Decide is an encoder that scores every permitted answer, then a constrained decoder picks the best joint set.
Skipping token-by-token generation is where the speed comes from. It is also why a decision model can ask several questions about the same input for little extra cost: the state is read once and every question is scored against it.
Three question types across vendors
Almost every product in the category uses the same three primitives, first set out in TypeSafe’s documentation. The names differ slightly, which matters if you plan to switch vendors later.
| Question type | What it returns | Name used by | Typical use |
|---|---|---|---|
| Yes/no | Probability from 0 to 1 that a condition is true | “noul” (TypeSafe, Cloudflare, Databricks, AWS); “predicate” (OpenAI) | Is this urgent? Is the tool call grounded? |
| Choice | One option plus a probability for each | “choice” (TypeSafe, Cloudflare, Databricks, AWS, OpenAI) | Which team, tool or queue? |
| Score | Probability-weighted average over ordered levels | “score” (TypeSafe, Cloudflare, Databricks, OpenAI) | How severe? How relevant? |
OpenAI’s documentation shows how a score works. With probabilities of 0.1, 0.7 and 0.2 across three severity levels numbered 0, 1 and 2, the returned score is 1.1, a value that falls between levels. That is useful for ranking, but a threshold on a score is a business decision, not a property of the model.
What a decision model cannot do
AWS is candid about the trade. Its launch post says the single parallel pass makes Strands Decider “significantly worse at solving complex problems than reasoning models”, and the lack of text output makes it unsuitable for coding, chatbots and document summaries. OpenAI’s guide says the same in practical terms: use Structured Outputs when you need generated fields or explanations, and function calling when a model must request a tool with arguments.
Inputs differ too. Jev reads text only, so images must be turned into text first. Clef reads text, JSON, images and video. OpenAI’s endpoint accepts images, but only as inline base64 data, not hosted URLs.
Where a Decision Model Sits in an Agent Stack
The practical question for an engineering team is which calls to move. A decision model fits wherever an agent pauses to choose from a known list. AWS lists the early wins it has seen: model routing, tool selection, evaluations, guardrails, memory, context management and policy classification.
Routing and next-step selection
The simplest use is the one that runs most often. An agent finishes a step and must pick the next tool, the next sub-agent or the next queue. Today many agents send that choice back to the same frontier model that does the planning, paying full price for a one-word answer. A hybrid design sends the routine choice to the decision model and keeps the expensive model for the hard cases, which AWS describes as “using LLMs to make the hardest decisions and using decider models to make the easier rote decisions”.
A guardrail in front of every tool call
The Strands Decider launch post includes a worked guardrail. Before an over-eager agent calls a weather tool, the decision model answers two yes/no questions. Are the tool’s arguments grounded in what the user actually said? Is it premature to call the tool before asking a clarifying question? A few lines of Python turn the probabilities into one of four typed actions: Proceed, Deny, Confirm (ask a person) or Guide (hand the turn back with feedback).
AWS’s own summary is the strongest argument for the layer: “a decision this cheap can sit in a path where an LLM call never could.” This matches the sandboxing work AWS described in Strands Box, where the goal is to stop a runaway agent before an action lands rather than after.
Triage, classification and moderation
Cloudflare’s own test came from its threat intelligence team, a cybersecurity unit that classifies websites. Given a domain, Clef fetched, rendered and classified the site in 2.2 seconds. Cloudflare’s fastest general model in the same workflow, gpt-oss-120b, took 4.7 seconds and returned only two categories. Moderation is a similar fit, as we found when covering AI content moderation with decision models.
| Agent job | Decision model | Reasoning LLM |
|---|---|---|
| Pick the next tool from a fixed list | Yes | Only when no option fits |
| Check a tool call before it runs | Yes | For a written explanation to a reviewer |
| Route and prioritise a ticket | Yes | To draft the reply |
| Plan a multi-step task | No | Yes |
| Write code or a summary | No | Yes |
Fastino adds a subtler point about guardrails. Asked two separate questions, its model flagged a prompt injection at 0.82 but also labelled the same prompt “safe” at 0.52. Joint decoding under a rule that any detected harm means “unsafe” removed the contradiction. If your decision model scores questions independently, your own code has to enforce that kind of consistency.
The Price and Latency Case for a Decision Model
Every vendor in the category prices well below frontier models, but the spread inside the category is still almost six times from cheapest to dearest:
All four charge for input only. TypeSafe says output tokens are free, and OpenAI says the Decisions API has no cache-read, cache-write or output-token charges, although regional processing premiums and long-context multipliers still apply. The open-weight options, Strands Decider 2B and GLiNER2.5-Decide, cost whatever your own hardware costs.
What a billion decision tokens costs
At list prices, one billion input tokens costs $42 on Jev, $90 on Clef-flash, $100 on OpenAI’s Decisions API and $240 on Clef. Those are small numbers next to frontier-model pricing, which is why InfoWorld’s sources see a decision model as a budget tool. Our guide to controlling AI token costs covers the wider set of levers.
The gap between the cheapest and dearest decision model is $198 per billion tokens. Hold that number, because later in this article the cost of a wrong decision dwarfs it.
Latency, measured by a competitor
Cloudflare published median latencies across 43 benchmarks for its own models and several rivals. The numbers come from one harness, which makes them comparable, but the harness belongs to a vendor with a product to sell:
Other vendors report on their own hardware. AWS measured a median of about 115 ms for Strands Decider on an Nvidia RTX 3090 and around 153 ms for small tasks on an M3 MacBook. Fastino reports a p50 of 38.3 ms on an Nvidia V100 and 167.3 ms on a 48-vCPU Intel Xeon. TypeSafe told InfoWorld that Jev answers in 70 to 500 ms. OpenAI says its endpoint is about ten times faster than the Responses API.
Accuracy is not settled
Every accuracy figure in the category is vendor-run. Cloudflare’s table shows Clef ahead of Jev on most benchmarks, such as BANKING77 (94.20 against 79.74), but Jev still wins When2Call (80.97 against 72.37) and BRIGHT. Fastino’s 60.1% average beats a model called JevK5, which Fastino itself notes is an open reproduction, not TypeSafe’s Jev. Treat each scoreboard as a reason to test, not a reason to buy.
AI-Stack Sprawl: The Hidden Cost of a Decision Model Layer
The second half of InfoWorld’s report is the warning. Ashish Chaturvedi, executive research leader at HFS Research, said: “As enterprises add these decision models alongside reasoning models, rerankers, embedding models, routers, guardrails and other specialized components, there is a significant risk of AI-stack sprawl.” He added that it will show up “in different places than people expect, most likely in evaluation and calibration”.
Every new model is another evaluation surface
Each model in the stack needs its own test set, its own monitoring, its own failure modes and its own upgrade plan. A team that runs one frontier model today might run a frontier model, a cheaper model, an embedding model, a reranker, a guardrail model and now a decision model tomorrow. Each one can drift, and each one can change underneath you when a vendor ships a new version.
That overhead is easy to miss in a pilot, because pilots run for weeks, not years. It shows up in month six, when an agent starts routing tickets to the wrong team and nobody can say whether the reasoning model, the decision model or the prompt changed.
Version drift you did not schedule
Several products in the category are built to move. TypeSafe’s model page says the jev-latest alias points at the newest stable release and is the default in its SDKs, while jev-preview can move ahead of it. The Databricks ai_decide documentation says Databricks “might change the model” if a better one emerges on its internal benchmarks. OpenAI expects its Decisions API to reach general availability “in the coming weeks”.
None of that is a flaw. It does mean a decision model needs the same version discipline as any other production dependency: pin the version, record it with every decision, and re-run your evaluation set before you move.
Why One Decision Model's 0.8 Is Not Another's
Calibration is the part of the decision model story that is easiest to underrate. A model is well calibrated when, of all the answers it gives at 0.8 confidence, about 80% are right. Chaturvedi’s warning to InfoWorld was specific: “A probability of 0.8 from Clef, Decider, and Luna will not reflect the same level of reliability, so thresholds tuned for one model will not transfer to another.”
He went on: “An enterprise that switches vendors, or runs several, has to recalibrate every threshold on its own data.” In other words, the threshold belongs to the pairing of a model and a decision, not to the decision alone.
How vendors train for calibration
The vendors do take calibration seriously. Cloudflare trained Clef with a label-smoothed cross-entropy loss plus a Brier loss “to refine probability calibration”, and added a reinforcement learning objective it calls Reinforcement Learning for Calibrated Decisions. AWS reports calibration for Strands Decider as a Brier score on JevBench’s public set and says the model ranked third of 33 in the 2B class.
Those are the vendors’ own data sets, though. OpenAI’s advice is the one to follow: “Use labeled examples from your application to set thresholds,” and choose them “based on the cost of false positives and false negatives.”
Confidence on questions with no right answer
Independent tests have found gaps. An outside phishing benchmark we covered in our piece on TypeSafe AI’s developer adoption found Jev scored 62.6% when asked one broad question but 95.0% when the same judgement was split into five narrow ones. It was also right only 44.7% of the time on unanswerable questions while reporting 0.74 confidence.
The lesson applies to every decision model, not just Jev. Narrow questions calibrate better than broad ones, and a “none of these” or “other” option gives the model somewhere honest to put uncertainty. OpenAI’s guide recommends exactly that fallback option.
A recalibration routine you can run
Start with a labelled set of real decisions for each schema, drawn from production traffic rather than written by the team. Plot how often the model is right at each confidence band. Set thresholds per model and per decision from that chart, weighted by what a false positive and a false negative cost you. Repeat the exercise whenever the model version, the question wording or the input mix changes.
Decision Schemas Are Business Policy
Chaturvedi’s third point is the one most likely to catch governance teams out. “Every question, answer set, and threshold encodes a piece of business policy, such as when a refund needs approval or what counts as a critical incident,” he told InfoWorld. “As teams create hundreds of these, they become a new body of logic that needs version control, ownership, and review, much as prompts did before them.”
What a decision schema registry should hold
A schema registry does not need new software. A version-controlled repository with a small amount of structure is enough. The minimum fields are below.
| Field | Why it matters | Usual owner |
|---|---|---|
| Question wording and answer set | Small wording changes move the scores | Business process owner |
| Threshold and the action it triggers | This is the policy itself | Business process owner |
| Pinned model and version | Thresholds do not transfer between models | Platform team |
| Evaluation set and last calibration date | Proves the threshold still holds | Platform team |
| Escalation path and human reviewer | Required where decisions affect people | Risk or compliance |
Stephanie Walter of HyperFrame Research made the same point in InfoWorld’s earlier coverage: developers have to specify questions, possible outputs, thresholds and escalation paths in advance, “which could be a significant task”. That work is policy design, so it needs policy owners, not just developers.
Conflicting decisions across teams
Without a registry, two teams will write two refund schemas with different thresholds, and the same customer will get two different answers depending on which agent picks up the case. A single owner per decision type, plus a rule that agents call the shared schema rather than writing their own, prevents most of it. Our IT governance work with clients starts with exactly this kind of inventory.
Audit trails and data handling
Paul Chada of Doozer AI told InfoWorld that regulated industries would ask about auditability. Log the input, the schema version, the model version, every probability returned, the threshold applied and the action taken. In the UK, decisions with legal or similarly significant effects on people also bring the ICO’s rules on automated decision-making into play. The NIST AI Risk Management Framework is a useful structure for the wider controls.
Data handling differs by vendor. Cloudflare says it does not read, store or train on Clef requests unless you use its fine-tuning service. OpenAI supports Zero Data Retention and HIPAA use for eligible customers, with processing in the US and Europe. Databricks processes data inside its own security perimeter and keeps run metadata.
Cost per Successful Decision, Not Cost per Token
Aditya Ranjan, a senior data engineer at the US supermarket chain H-E-B, gave InfoWorld the most practical warning. “A decision model might reduce inference spending significantly, but if the enterprise needs additional engineering resources to maintain multiple models, manage failures, and investigate inconsistent results, the financial benefit may be smaller than expected,” he said.
The error bill is bigger than the inference bill
Ranjan added that a wrong decision can trigger a downstream action that costs more to fix than the inference ever did. A simple example shows the scale. These assumptions are illustrative, not drawn from any vendor: 2 million decisions a day, 1,500 input tokens each, so 3 billion tokens a day, and $3 to unwind each wrong decision with a human review and a reversal.
At list prices, that volume costs $126 a day on Jev, $270 on Clef-flash, $300 on the Decisions API and $720 on Clef. At a 1% error rate, 20,000 wrong decisions a day cost $60,000 to unwind. At 0.5%, they cost $30,000:
Cutting the error rate from 1.0% to 0.8% in this example saves $12,000 a day. That is about 17 times the whole daily inference bill on the most expensive decision model. Pick the decision model that makes the fewest costly mistakes on your data, then worry about the price.
The five metrics to track
Ranjan’s recommended scorecard for CIOs has five measures: end-to-end workflow latency, cost per successful decision, decision accuracy, operational overhead and failure recovery costs. Cost per successful decision is the one that joins them up. Divide everything the decision layer costs, including people and error clean-up, by the number of decisions that were right and did not need a human to fix them.
Build, Rent or Abstract the Decision Model Layer
InfoWorld’s report ends on a hopeful note for CIOs. OpenAI says its Decisions API could abstract away some of the complexity, because developers define the decision and the answers and never pick, deploy or manage a separate model. We looked at the launch in detail in our article on OpenAI’s Decisions API.
InfoWorld is also clear about the limit: “the abstraction does not eliminate the need for enterprises to evaluate and govern the decisions being made.” The registry, the thresholds and the audit trail stay with you whichever route you take.
Rent a hosted decision model
Jev, Clef on Workers AI and the OpenAI Decisions API all run as hosted services. That is the fastest route and keeps the infrastructure off your books. The trade is that the vendor controls versions, rate limits and prices. TypeSafe currently caps Jev at 100,000 tokens and 80 requests per second.
Run open weights you control
Strands Decider 2B is open source with all of its training data and scripts, and runs on a local CPU or GPU. GLiNER2.5-Decide is Apache 2.0, runs on CPUs and can be deployed in air-gapped environments. Cloudflare has also published Clef’s weights on Hugging Face under Apache 2.0, though the 27B model needs serious GPU memory. Open weights keep a record of every routing decision inside your own network, which matters to regulated buyers.
Fine-tune for your own decisions
The options here differ most. Cloudflare is offering fine-tuning first through its forward-deployed engineers and later as a self-serve reinforcement learning platform built on AI Gateway, Workers AI and Containers. Fastino supports full and LoRA fine-tuning. TypeSafe’s documentation steers customers to put their own rules into the state, the question instructions and the criteria instead.
| Route | Examples | You still manage | Best when |
|---|---|---|---|
| Abstract | OpenAI Decisions API, Databricks ai_decide | Schemas, thresholds, audit | You already use that platform |
| Rent a named model | Jev, Clef and Clef-flash on Workers AI | Version pinning and calibration | You want to choose and compare models |
| Run open weights | Strands Decider 2B, GLiNER2.5-Decide, Clef weights | Hosting, scaling, upgrades | Data must stay in your network |
A 90-Day Plan for Adding a Decision Model Layer
The safest way in is to treat the decision layer as an optimisation of decisions your agents already make, not as a new capability. This plan fits a team with at least one agent in production. It also fits neatly into a wider AI strategy review.
Days 1 to 30: inventory the decisions
List every point where an agent or workflow chooses from a fixed set: routing, tool choice, approval checks, triage, moderation. For each, record the volume per day, the current model, the cost and what a wrong answer costs. Rank them by volume multiplied by error cost. The top five are your candidates.
Days 31 to 60: run in shadow mode
Run one or two decision models alongside your current setup on the candidate decisions, without acting on their answers. Compare their choices with the current model’s and with human labels. Build the calibration chart for each model and decision, and set thresholds from it. Record everything in the schema registry from day one.
Days 61 to 90: cut over behind gates
Switch the best-performing decision to the new model with a pinned version, a fallback “other” answer routed to a person, and an alert if the share of low-confidence answers rises. Review cost per successful decision weekly. Only then move the next decision across.
Questions Buyers Are Asking About the Decision Model Layer
Is a decision model just a classifier?
Partly. Classifiers have been around for years, but they had to be retrained for each new set of categories. Cloudflare’s point is that a decision model works over any set of inputs and options you define at request time, “without constantly retraining the model to incorporate new classification categories”.
Will a decision model replace our LLM?
No. Every vendor positions it alongside a reasoning model, not instead of one. The decision model handles the frequent, bounded choices, and the LLM keeps planning, writing and anything that needs open-ended reasoning.
Can we switch decision model vendors later?
Technically yes, since Clef is “fully Jev-API compatible” and most products share the same three question types. Operationally it costs a full recalibration, because thresholds do not transfer. Keep schemas vendor-neutral and budget for re-testing.
Does the decision model layer need its own budget line?
It should. The inference bill is small, but evaluation, calibration, schema governance and error handling are real costs. Putting them on one line is the only way to see the cost per successful decision that Ranjan recommends.
References
As AI agents grow, vendors carve out decision-making as a separate model layer (InfoWorld)
TypeSafe AI’s new models work with machines, not humans (InfoWorld)
Introducing Clef: open-source decision models and RL fine-tuning (Cloudflare)
Clef model documentation (Cloudflare Workers AI)
Clef-flash model documentation (Cloudflare Workers AI)
Introducing Strands Decider 2B (Strands Agents)
strands-decider repository (GitHub)
Jev models, pricing and limits (TypeSafe docs)
Introducing System One models and Jev (TypeSafe)
GLiNER2.5-Decide: an open-weight model for structured decision making (Fastino)
ai_decide function (Databricks documentation)
Cloudflare Clef model card (Hugging Face)
AI Risk Management Framework (NIST)
Rights related to automated decision-making including profiling (ICO)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.