Decision model releases from AWS, Cloudflare and OpenAI arrived within days of each other at the start of October, and on 9 October InfoWorld reported that the pattern now looks like a new tier of the enterprise AI stack. The idea is simple. The bounded choices an agent makes all day go to a small model built only to choose. Which tool should it call? Which team gets this ticket? Does this action need a human? A large reasoning model keeps the planning, the coding and the writing.

That split promises smaller token bills and faster AI agents. The analysts InfoWorld spoke to also warn of a new kind of AI-stack sprawl. Enterprises get more models to evaluate, confidence scores that mean different things from vendor to vendor, and hundreds of decision schemas that encode business policy without a clear owner.

This article sets out what each vendor shipped, how a decision model differs from the large language model above it, and what the new layer costs and saves. It then covers the calibration and governance work that decides whether the savings survive contact with production. It follows our coverage of the Jev model launch in September, and focuses on what InfoWorld’s analysts say the new layer costs to run.

What InfoWorld Reported About the Decision Model Layer

decision model layer ai agents stack sprawl b weather house with one figure stepped out

InfoWorld’s Anirban Ghoshal framed the trend around cost. Enterprises are trying to balance AI budgets against the demands of scaling agentic applications. They have tried hard-coding deterministic decisions and using smaller models for specific tasks. TypeSafe’s Jev, released in mid-September, added a third option: a specialised decision model for the bounded decisions that sit between an agent’s reasoning and its actions.

In the last week of September and the first days of October, three large vendors followed. Cloudflare released Clef and Clef-flash. AWS released Strands Decider 2B. OpenAI released a Decisions API that hides the model behind an endpoint. InfoWorld’s conclusion was that decision-making “could become another specialized layer in the enterprise AI stack”.

Six decision model launches in under three weeks

The table below lists what shipped, using each vendor’s own documentation. TypeSafe and Fastino had already started the category by the time the hyperscalers arrived, and Databricks added a SQL function in the same week as OpenAI. Our earlier piece on the wave of Jev alternatives covers the smaller open-source copies in more detail.

Vendor and modelReleasedSize and licenceList price (input)Inputs
TypeSafe Jev 1.1315 Sep 2026Size not published; hosted API$0.042 per million tokens; output freeText only
Fastino GLiNER2.5-Decide24 Sep 2026340M encoder; Apache 2.0Open weights; your hardwareText
Databricks ai_decideEnd of Sep 2026 (beta)Apache 2.0 models chosen by DatabricksDatabricks SQL pricingText or structured data
OpenAI Decisions API30 Sep 2026 (public beta)gpt-6-luna; model hidden$0.10 per million tokens; no output chargeText and images
Cloudflare Clef / Clef-flash1 Oct 202627B / 9B; Apache 2.0 weights$0.24 / $0.09 per million tokensText, JSON, images, video
AWS Strands Decider 2B1 Oct 20262B; open source, data and scripts includedOpen weights; your hardwareText

Why InfoWorld calls it a layer

The word “layer” is doing real work in the headline. AWS describes Strands Decider as a control-flow component that picks tools, routes tasks and decides an agent’s next step, which InfoWorld summarised as “in effect acting as an orchestrator for what an agent does next”. Cloudflare’s launch post suggests putting Clef “into the hot path for agents to make decisions” and pairing it with a larger model on Workers AI to take the action.

So the decision model is not a cheaper chatbot. It sits between reasoning and action and turns a fuzzy situation into a typed answer that ordinary code can branch on. David Linthicum, quoted in InfoWorld’s earlier coverage of Jev, put the old approach bluntly: using a general-purpose LLM for every bounded choice “is like using a full enterprise service bus to answer a yes/no routing question”.

How a Decision Model Differs From the LLM Above It

decision model layer ai agents stack sprawl c wishbone standing in a round display stand

A decision model does not write. It reads a state (a support ticket, an agent transcript, a product photo) plus a set of typed questions, and returns a probability for every allowed answer. Because there is no free text to parse, the output is always one of the options you defined.

One pass, no generated text

The vendors use different architectures to get the same effect. Cloudflare runs its Qwen backbone in a single prefill-only pass and then scores the valid schema choices in parallel, so “there’s no intermediate text to generate token by token”. AWS takes a Qwen3.5-2B model, removes the head that generates text and replaces it with a pointer head of just over a million parameters that scores each option. Fastino’s GLiNER2.5-Decide is an encoder that scores every permitted answer, then a constrained decoder picks the best joint set.

Skipping token-by-token generation is where the speed comes from. It is also why a decision model can ask several questions about the same input for little extra cost: the state is read once and every question is scored against it.

Three question types across vendors

Almost every product in the category uses the same three primitives, first set out in TypeSafe’s documentation. The names differ slightly, which matters if you plan to switch vendors later.

Question typeWhat it returnsName used byTypical use
Yes/noProbability from 0 to 1 that a condition is true“noul” (TypeSafe, Cloudflare, Databricks, AWS); “predicate” (OpenAI)Is this urgent? Is the tool call grounded?
ChoiceOne option plus a probability for each“choice” (TypeSafe, Cloudflare, Databricks, AWS, OpenAI)Which team, tool or queue?
ScoreProbability-weighted average over ordered levels“score” (TypeSafe, Cloudflare, Databricks, OpenAI)How severe? How relevant?

OpenAI’s documentation shows how a score works. With probabilities of 0.1, 0.7 and 0.2 across three severity levels numbered 0, 1 and 2, the returned score is 1.1, a value that falls between levels. That is useful for ranking, but a threshold on a score is a business decision, not a property of the model.

What a decision model cannot do

AWS is candid about the trade. Its launch post says the single parallel pass makes Strands Decider “significantly worse at solving complex problems than reasoning models”, and the lack of text output makes it unsuitable for coding, chatbots and document summaries. OpenAI’s guide says the same in practical terms: use Structured Outputs when you need generated fields or explanations, and function calling when a model must request a tool with arguments.

Inputs differ too. Jev reads text only, so images must be turned into text first. Clef reads text, JSON, images and video. OpenAI’s endpoint accepts images, but only as inline base64 data, not hosted URLs.

Where a Decision Model Sits in an Agent Stack

decision model layer ai agents stack sprawl d garden hose reel cart with the hose coiled on its drum

The practical question for an engineering team is which calls to move. A decision model fits wherever an agent pauses to choose from a known list. AWS lists the early wins it has seen: model routing, tool selection, evaluations, guardrails, memory, context management and policy classification.

Routing and next-step selection

The simplest use is the one that runs most often. An agent finishes a step and must pick the next tool, the next sub-agent or the next queue. Today many agents send that choice back to the same frontier model that does the planning, paying full price for a one-word answer. A hybrid design sends the routine choice to the decision model and keeps the expensive model for the hard cases, which AWS describes as “using LLMs to make the hardest decisions and using decider models to make the easier rote decisions”.

A guardrail in front of every tool call

The Strands Decider launch post includes a worked guardrail. Before an over-eager agent calls a weather tool, the decision model answers two yes/no questions. Are the tool’s arguments grounded in what the user actually said? Is it premature to call the tool before asking a clarifying question? A few lines of Python turn the probabilities into one of four typed actions: Proceed, Deny, Confirm (ask a person) or Guide (hand the turn back with feedback).

AWS’s own summary is the strongest argument for the layer: “a decision this cheap can sit in a path where an LLM call never could.” This matches the sandboxing work AWS described in Strands Box, where the goal is to stop a runaway agent before an action lands rather than after.

Triage, classification and moderation

Cloudflare’s own test came from its threat intelligence team, a cybersecurity unit that classifies websites. Given a domain, Clef fetched, rendered and classified the site in 2.2 seconds. Cloudflare’s fastest general model in the same workflow, gpt-oss-120b, took 4.7 seconds and returned only two categories. Moderation is a similar fit, as we found when covering AI content moderation with decision models.

Agent jobDecision modelReasoning LLM
Pick the next tool from a fixed listYesOnly when no option fits
Check a tool call before it runsYesFor a written explanation to a reviewer
Route and prioritise a ticketYesTo draft the reply
Plan a multi-step taskNoYes
Write code or a summaryNoYes

Fastino adds a subtler point about guardrails. Asked two separate questions, its model flagged a prompt injection at 0.82 but also labelled the same prompt “safe” at 0.52. Joint decoding under a rule that any detected harm means “unsafe” removed the contradiction. If your decision model scores questions independently, your own code has to enforce that kind of consistency.

The Price and Latency Case for a Decision Model

decision model layer ai agents stack sprawl e fidget spinner spinning on a desk stand

Every vendor in the category prices well below frontier models, but the spread inside the category is still almost six times from cheapest to dearest:

List price per million input tokens (US dollars)
TypeSafe Jev 1.13 $0.042
Cloudflare Clef-flash $0.09
OpenAI Decisions API (gpt-6-luna) $0.10
Cloudflare Clef $0.24

All four charge for input only. TypeSafe says output tokens are free, and OpenAI says the Decisions API has no cache-read, cache-write or output-token charges, although regional processing premiums and long-context multipliers still apply. The open-weight options, Strands Decider 2B and GLiNER2.5-Decide, cost whatever your own hardware costs.

What a billion decision tokens costs

At list prices, one billion input tokens costs $42 on Jev, $90 on Clef-flash, $100 on OpenAI’s Decisions API and $240 on Clef. Those are small numbers next to frontier-model pricing, which is why InfoWorld’s sources see a decision model as a budget tool. Our guide to controlling AI token costs covers the wider set of levers.

The gap between the cheapest and dearest decision model is $198 per billion tokens. Hold that number, because later in this article the cost of a wrong decision dwarfs it.

Latency, measured by a competitor

Cloudflare published median latencies across 43 benchmarks for its own models and several rivals. The numbers come from one harness, which makes them comparable, but the harness belongs to a vendor with a product to sell:

Median decision latency on Cloudflare’s benchmark (milliseconds)
Laya 5.8 ms
Clef-flash 38.8 ms
Kev-9B 51.4 ms
Clef 209.3 ms
Jev 524.1 ms

Other vendors report on their own hardware. AWS measured a median of about 115 ms for Strands Decider on an Nvidia RTX 3090 and around 153 ms for small tasks on an M3 MacBook. Fastino reports a p50 of 38.3 ms on an Nvidia V100 and 167.3 ms on a 48-vCPU Intel Xeon. TypeSafe told InfoWorld that Jev answers in 70 to 500 ms. OpenAI says its endpoint is about ten times faster than the Responses API.

Accuracy is not settled

Every accuracy figure in the category is vendor-run. Cloudflare’s table shows Clef ahead of Jev on most benchmarks, such as BANKING77 (94.20 against 79.74), but Jev still wins When2Call (80.97 against 72.37) and BRIGHT. Fastino’s 60.1% average beats a model called JevK5, which Fastino itself notes is an open reproduction, not TypeSafe’s Jev. Treat each scoreboard as a reason to test, not a reason to buy.

AI-Stack Sprawl: The Hidden Cost of a Decision Model Layer

decision model layer ai agents stack sprawl f banyan tree spreading on its prop roots

The second half of InfoWorld’s report is the warning. Ashish Chaturvedi, executive research leader at HFS Research, said: “As enterprises add these decision models alongside reasoning models, rerankers, embedding models, routers, guardrails and other specialized components, there is a significant risk of AI-stack sprawl.” He added that it will show up “in different places than people expect, most likely in evaluation and calibration”.

Every new model is another evaluation surface

Each model in the stack needs its own test set, its own monitoring, its own failure modes and its own upgrade plan. A team that runs one frontier model today might run a frontier model, a cheaper model, an embedding model, a reranker, a guardrail model and now a decision model tomorrow. Each one can drift, and each one can change underneath you when a vendor ships a new version.

That overhead is easy to miss in a pilot, because pilots run for weeks, not years. It shows up in month six, when an agent starts routing tickets to the wrong team and nobody can say whether the reasoning model, the decision model or the prompt changed.

Version drift you did not schedule

Several products in the category are built to move. TypeSafe’s model page says the jev-latest alias points at the newest stable release and is the default in its SDKs, while jev-preview can move ahead of it. The Databricks ai_decide documentation says Databricks “might change the model” if a better one emerges on its internal benchmarks. OpenAI expects its Decisions API to reach general availability “in the coming weeks”.

None of that is a flaw. It does mean a decision model needs the same version discipline as any other production dependency: pin the version, record it with every decision, and re-run your evaluation set before you move.

Why One Decision Model's 0.8 Is Not Another's

Calibration is the part of the decision model story that is easiest to underrate. A model is well calibrated when, of all the answers it gives at 0.8 confidence, about 80% are right. Chaturvedi’s warning to InfoWorld was specific: “A probability of 0.8 from Clef, Decider, and Luna will not reflect the same level of reliability, so thresholds tuned for one model will not transfer to another.”

He went on: “An enterprise that switches vendors, or runs several, has to recalibrate every threshold on its own data.” In other words, the threshold belongs to the pairing of a model and a decision, not to the decision alone.

How vendors train for calibration

The vendors do take calibration seriously. Cloudflare trained Clef with a label-smoothed cross-entropy loss plus a Brier loss “to refine probability calibration”, and added a reinforcement learning objective it calls Reinforcement Learning for Calibrated Decisions. AWS reports calibration for Strands Decider as a Brier score on JevBench’s public set and says the model ranked third of 33 in the 2B class.

Those are the vendors’ own data sets, though. OpenAI’s advice is the one to follow: “Use labeled examples from your application to set thresholds,” and choose them “based on the cost of false positives and false negatives.”

Confidence on questions with no right answer

Independent tests have found gaps. An outside phishing benchmark we covered in our piece on TypeSafe AI’s developer adoption found Jev scored 62.6% when asked one broad question but 95.0% when the same judgement was split into five narrow ones. It was also right only 44.7% of the time on unanswerable questions while reporting 0.74 confidence.

The lesson applies to every decision model, not just Jev. Narrow questions calibrate better than broad ones, and a “none of these” or “other” option gives the model somewhere honest to put uncertainty. OpenAI’s guide recommends exactly that fallback option.

A recalibration routine you can run

Start with a labelled set of real decisions for each schema, drawn from production traffic rather than written by the team. Plot how often the model is right at each confidence band. Set thresholds per model and per decision from that chart, weighted by what a false positive and a false negative cost you. Repeat the exercise whenever the model version, the question wording or the input mix changes.

Decision Schemas Are Business Policy

Chaturvedi’s third point is the one most likely to catch governance teams out. “Every question, answer set, and threshold encodes a piece of business policy, such as when a refund needs approval or what counts as a critical incident,” he told InfoWorld. “As teams create hundreds of these, they become a new body of logic that needs version control, ownership, and review, much as prompts did before them.”

What a decision schema registry should hold

A schema registry does not need new software. A version-controlled repository with a small amount of structure is enough. The minimum fields are below.

FieldWhy it mattersUsual owner
Question wording and answer setSmall wording changes move the scoresBusiness process owner
Threshold and the action it triggersThis is the policy itselfBusiness process owner
Pinned model and versionThresholds do not transfer between modelsPlatform team
Evaluation set and last calibration dateProves the threshold still holdsPlatform team
Escalation path and human reviewerRequired where decisions affect peopleRisk or compliance

Stephanie Walter of HyperFrame Research made the same point in InfoWorld’s earlier coverage: developers have to specify questions, possible outputs, thresholds and escalation paths in advance, “which could be a significant task”. That work is policy design, so it needs policy owners, not just developers.

Conflicting decisions across teams

Without a registry, two teams will write two refund schemas with different thresholds, and the same customer will get two different answers depending on which agent picks up the case. A single owner per decision type, plus a rule that agents call the shared schema rather than writing their own, prevents most of it. Our IT governance work with clients starts with exactly this kind of inventory.

Audit trails and data handling

Paul Chada of Doozer AI told InfoWorld that regulated industries would ask about auditability. Log the input, the schema version, the model version, every probability returned, the threshold applied and the action taken. In the UK, decisions with legal or similarly significant effects on people also bring the ICO’s rules on automated decision-making into play. The NIST AI Risk Management Framework is a useful structure for the wider controls.

Data handling differs by vendor. Cloudflare says it does not read, store or train on Clef requests unless you use its fine-tuning service. OpenAI supports Zero Data Retention and HIPAA use for eligible customers, with processing in the US and Europe. Databricks processes data inside its own security perimeter and keeps run metadata.

Cost per Successful Decision, Not Cost per Token

Aditya Ranjan, a senior data engineer at the US supermarket chain H-E-B, gave InfoWorld the most practical warning. “A decision model might reduce inference spending significantly, but if the enterprise needs additional engineering resources to maintain multiple models, manage failures, and investigate inconsistent results, the financial benefit may be smaller than expected,” he said.

The error bill is bigger than the inference bill

Ranjan added that a wrong decision can trigger a downstream action that costs more to fix than the inference ever did. A simple example shows the scale. These assumptions are illustrative, not drawn from any vendor: 2 million decisions a day, 1,500 input tokens each, so 3 billion tokens a day, and $3 to unwind each wrong decision with a human review and a reversal.

At list prices, that volume costs $126 a day on Jev, $270 on Clef-flash, $300 on the Decisions API and $720 on Clef. At a 1% error rate, 20,000 wrong decisions a day cost $60,000 to unwind. At 0.5%, they cost $30,000:

Illustrative daily cost at 2 million decisions (US dollars)
Errors at a 1% error rate $60,000
Errors at a 0.5% error rate $30,000
Inference on Clef $720
Inference on the Decisions API $300
Inference on Jev $126

Cutting the error rate from 1.0% to 0.8% in this example saves $12,000 a day. That is about 17 times the whole daily inference bill on the most expensive decision model. Pick the decision model that makes the fewest costly mistakes on your data, then worry about the price.

The five metrics to track

Ranjan’s recommended scorecard for CIOs has five measures: end-to-end workflow latency, cost per successful decision, decision accuracy, operational overhead and failure recovery costs. Cost per successful decision is the one that joins them up. Divide everything the decision layer costs, including people and error clean-up, by the number of decisions that were right and did not need a human to fix them.

Build, Rent or Abstract the Decision Model Layer

InfoWorld’s report ends on a hopeful note for CIOs. OpenAI says its Decisions API could abstract away some of the complexity, because developers define the decision and the answers and never pick, deploy or manage a separate model. We looked at the launch in detail in our article on OpenAI’s Decisions API.

InfoWorld is also clear about the limit: “the abstraction does not eliminate the need for enterprises to evaluate and govern the decisions being made.” The registry, the thresholds and the audit trail stay with you whichever route you take.

Rent a hosted decision model

Jev, Clef on Workers AI and the OpenAI Decisions API all run as hosted services. That is the fastest route and keeps the infrastructure off your books. The trade is that the vendor controls versions, rate limits and prices. TypeSafe currently caps Jev at 100,000 tokens and 80 requests per second.

Run open weights you control

Strands Decider 2B is open source with all of its training data and scripts, and runs on a local CPU or GPU. GLiNER2.5-Decide is Apache 2.0, runs on CPUs and can be deployed in air-gapped environments. Cloudflare has also published Clef’s weights on Hugging Face under Apache 2.0, though the 27B model needs serious GPU memory. Open weights keep a record of every routing decision inside your own network, which matters to regulated buyers.

Fine-tune for your own decisions

The options here differ most. Cloudflare is offering fine-tuning first through its forward-deployed engineers and later as a self-serve reinforcement learning platform built on AI Gateway, Workers AI and Containers. Fastino supports full and LoRA fine-tuning. TypeSafe’s documentation steers customers to put their own rules into the state, the question instructions and the criteria instead.

RouteExamplesYou still manageBest when
AbstractOpenAI Decisions API, Databricks ai_decideSchemas, thresholds, auditYou already use that platform
Rent a named modelJev, Clef and Clef-flash on Workers AIVersion pinning and calibrationYou want to choose and compare models
Run open weightsStrands Decider 2B, GLiNER2.5-Decide, Clef weightsHosting, scaling, upgradesData must stay in your network

A 90-Day Plan for Adding a Decision Model Layer

The safest way in is to treat the decision layer as an optimisation of decisions your agents already make, not as a new capability. This plan fits a team with at least one agent in production. It also fits neatly into a wider AI strategy review.

Days 1 to 30: inventory the decisions

List every point where an agent or workflow chooses from a fixed set: routing, tool choice, approval checks, triage, moderation. For each, record the volume per day, the current model, the cost and what a wrong answer costs. Rank them by volume multiplied by error cost. The top five are your candidates.

Days 31 to 60: run in shadow mode

Run one or two decision models alongside your current setup on the candidate decisions, without acting on their answers. Compare their choices with the current model’s and with human labels. Build the calibration chart for each model and decision, and set thresholds from it. Record everything in the schema registry from day one.

Days 61 to 90: cut over behind gates

Switch the best-performing decision to the new model with a pinned version, a fallback “other” answer routed to a person, and an alert if the share of low-confidence answers rises. Review cost per successful decision weekly. Only then move the next decision across.

Questions Buyers Are Asking About the Decision Model Layer

Is a decision model just a classifier?

Partly. Classifiers have been around for years, but they had to be retrained for each new set of categories. Cloudflare’s point is that a decision model works over any set of inputs and options you define at request time, “without constantly retraining the model to incorporate new classification categories”.

Will a decision model replace our LLM?

No. Every vendor positions it alongside a reasoning model, not instead of one. The decision model handles the frequent, bounded choices, and the LLM keeps planning, writing and anything that needs open-ended reasoning.

Can we switch decision model vendors later?

Technically yes, since Clef is “fully Jev-API compatible” and most products share the same three question types. Operationally it costs a full recalibration, because thresholds do not transfer. Keep schemas vendor-neutral and budget for re-testing.

Does the decision model layer need its own budget line?

It should. The inference bill is small, but evaluation, calibration, schema governance and error handling are real costs. Putting them on one line is the only way to see the cost per successful decision that Ranjan recommends.

References