Jev alternatives arrived faster than almost anyone expected. Within 16 days of TypeSafe AI releasing Jev on 15 September 2026, Amazon, Cloudflare, OpenAI and Databricks had each shipped a product that does the same job, and developers had published dozens of open-source copies. The Wall Street Journal reported on 2 October that the model “has set Silicon Valley circles abuzz with discussions around alternatives to large language models, and is already sparking copycats.”
Jev does not write text. It reads some input, answers questions with a fixed set of possible answers, and attaches a probability to every option. TypeSafe chief executive Diogo Almeida told the Journal that about 25% of Fortune 500 companies already use it, and that the model was handling a trillion tokens a day. Investors are reportedly discussing a valuation of $10 billion or more.
This article covers what the Journal reported, every major group of Jev alternatives released so far, how they compare on price, speed and accuracy, and whether decision models really are an alternative to large language models. For background, see our Jev model launch analysis, our report on how developers adopted TypeSafe AI, and our look at OpenAI’s Decisions API.
Table of contents
- What the Wall Street Journal Reported About Jev Alternatives
- The Big-Company Jev Alternatives, Week by Week
- The Open-Source Jev Alternatives Developers Are Starring
- How the Jev Alternatives Compare on Speed and Price
- Benchmarks: Where Jev Still Beats the Jev Alternatives
- Are Jev and the Jev Alternatives Really a Replacement for LLMs?
- What the Jev Alternatives Mean for TypeSafe
- How to Choose Between Jev Alternatives
- What to Watch Next for Jev Alternatives
- Jev Alternatives FAQ
- References and Further Reading
What the Wall Street Journal Reported About Jev Alternatives
The Journal’s story, by Elias Schisgall, is short. Its news value lies in three things: new adoption figures from Almeida, confirmation that the copycats now include large companies, and Almeida’s argument that decision models are the start of a new class of AI.
The adoption figures Almeida gave
Almeida spoke to the Journal on Monday 28 September. He said “some 25%” of Fortune 500 companies use Jev, and gave a throughput figure: “We were at a trillion tokens per day about a week ago. And obviously, growth has been exponential.”
Neither number has been independently checked, and TypeSafe has not said how it counts a company as a user, for example whether one developer testing the API is enough. Still, the figures fit what outside platforms reported. Vercel said Jev drew more than twice as much interest from paid developer accounts in its first day as any earlier model launch, and the Financial Times reported that tokens sent to Jev through OpenRouter more than tripled over one weekend.
The funding talks
On 24 September, The Information reported that TypeSafe had started talking to investors about raising “$1 billion or even more”, with some unnamed investors offering to invest at a valuation of $10 billion or more. The startup had announced only the week before that it had raised $40 million at a valuation of $200 million, according to PitchBook, in a round led by DCVC.
That would be a 50-fold rise on the valuation TypeSafe disclosed at launch. Almeida declined to comment on the reported round to the Journal. He told the Financial Times that potential investors were “battering down our door”. Most of the Jev alternatives described below arrived within a week of those reports.
Almeida’s case for leaving chatbots behind
The Journal’s headline phrase, “talk of LLM alternatives”, comes from Almeida’s argument. He said today’s leading models work well when a person is involved but are “unbelievably useless for automation”.
“Software tends to work by building a stable layer that people can build on top of, and layering on top of that, and repeating and repeating and repeating,” he said. “That has absolutely not happened with chatbots because they don’t behave that way.” He called Jev the start of a “new class of AI” and added: “There’s a whole bunch of frontiers, actually. It just so happens TypeSafe is making the first next frontier, and there will be many more.”
The Big-Company Jev Alternatives, Week by Week
The copycats came in a rush during the week of 28 September. AutoTrust AI launched first, then OpenAI, Databricks, Amazon and Cloudflare followed within four days, alongside Fastino Labs. The table lists each one, using the company’s own description.
| Date | Company and product | How it is built | How you get it |
|---|---|---|---|
| 28 Sep | AutoTrust AI, JEV-27B | Small decision block on a frozen Qwen3.8-27B | Open weights, Apache 2.0 |
| 29 Sep | OpenAI, Decisions API | A version of GPT-6 Luna | Hosted API, limited preview |
| 30 Sep | Databricks, ai_decide | An unnamed decision model | SQL function and REST API, beta |
| 1 Oct | Amazon, Strands Decider 2B | Qwen3.5-2B with its text head replaced | Open source, with weights and data |
| 1 Oct | Cloudflare, Clef and Clef-flash | Frozen Qwen3.8-27B and Qwen3.5-9B | Workers AI, or open weights |
| 1 Oct | Fastino Labs, GLiDE | Decision model that can stop and reason | Hosted API |
OpenAI and Databricks build Jev alternatives into their platforms
OpenAI announced its Decisions API at its DevDay conference on Tuesday 29 September. The Journal says it uses OpenAI’s Luna model to answer “a specific set of user-defined questions with finite pre-defined answers”. OpenAI has said it returns results in about 150 milliseconds and accepts images, but it has not published a price. Almeida’s reply on X was a joke: “begun, the clone war has”.
Databricks followed on 30 September with ai_decide, a function that runs inside its data platform. Its launch post names TypeSafe directly and says the function is “directly compatible” with the TypeSafe API. A Databricks engineer wrote on X: “Following the TypeSafe AI Jev launch, we’ve seen increasing demand for a fast, low-cost API that turns raw text into structured decisions.”
Amazon’s model began as one engineer’s side project
Amazon’s entry, Strands Decider 2B, started at home. Marc Brooker, a distinguished engineer at Amazon Web Services, saw Jev and set out to build his own small version on a home graphics card. He called it Hobson. It did well enough that, according to TechCrunch, it briefly topped the community JevBench ranking for its size, and AWS engineers tidied it up and released it through Strands Labs.
The design is easy to follow. Brooker took Qwen3.5-2B, removed the part that generates text, and added a small “pointer head” of just over a million parameters that scores each answer option. His blog says he used self-distillation during training to limit catastrophic forgetting, the tendency of a fine-tuned model to lose skills it already had. Amazon released the weights, the training data and the scripts.
“What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step,” Brooker told TechCrunch. He added that the cost of building something interesting in this space runs to “the hundreds or thousands of dollars”, which is why he does not expect the large labs to dominate it.
Cloudflare’s Clef adds images and costs more
Cloudflare announced Clef and the smaller Clef-flash on 1 October. Both use a frozen Qwen model that reads the input once, then score every allowed answer in parallel. Clef accepts images and video, which Jev cannot, and Cloudflare says its own tests rank Clef first on the community Jev Decision Index.
The price is higher. Clef costs $0.24 per million tokens on Workers AI, about 5.7 times Jev’s $0.042, by our arithmetic. The weights are free to download under Apache 2.0, but Cloudflare told The Register that its training data is not public. Running Clef yourself needs a GPU with 85 GB of memory, or 41 GB for Clef-flash.
Two smaller labs say their Jev alternatives beat Jev
AutoTrust AI, based in Singapore, released JEV-27B on 28 September. It trained a 108.9-million-parameter decision block on top of a frozen Qwen3.8-27B in about 9.2 hours on one Nvidia B200, and reports a six-benchmark average of 84.07% against 83.85% for Jev. Its press release says plainly that it ran the comparison itself, so the result is not independent validation.
Fastino Labs, maker of the GLiNER models, released GLiDE on 1 October. It calls GLiDE the first “thinking” decision model: it answers easy questions in one fast pass and reasons further when it is unsure. Fastino says GLiDE scores 64.81 on the Decision Index against Jev’s published 57.91, using the official scorer on its own runs.
The Open-Source Jev Alternatives Developers Are Starring
Outside the big companies, the number of open-source Jev alternatives is harder to measure. TechCrunch says “dozens” of similar models have appeared. A GitHub search we ran on 3 October found 11,894 repositories with “jev” in the name or description created since 15 September, against 3,468 created in all the years before. Not every match is a decision model, but the jump is hard to miss.
GitHub repositories with “jev” in the name or description (GitHub search, 3 October 2026; our arithmetic: 11,894 / 3,468 = 3.4 times)
The open-source Jev alternatives with the most attention
A handful of these Jev alternatives have attracted most of the interest. The star counts below were read from GitHub on 3 October. Stars measure attention, not quality.
| Project | Built on | Stars | Licence | What stands out |
|---|---|---|---|---|
| Laya | ModernBERT-large, 421M parameters | 30,404 | Apache 2.0 | 39.5 ms per question on a T4 GPU |
| Kev | Qwen3.5 and Qwen3.8, 0.8B to 27B | 8,358 | Apache 2.0 | Works with TypeSafe’s own SDK unchanged |
| SemIf (formerly OpenJev) | Any frozen open model | 4,673 | MIT | No training: reads answer probabilities directly |
| NanoJev | Qwen3-0.6B | 2,482 | MIT | Built for fast game and control loops |
| jevlike | Your own encoder | 1,341 | MIT | A training recipe, not a finished model |
| AnyJev (Nokia) | Any open model | 1,019 | Apache 2.0 | Turns an existing model into a decision model |
Copying the interface took about a day
The first open Jev alternatives appeared within hours. By 19 September the Latent Space newsletter had counted six clones in two days. The speed tells you something about what TypeSafe published: the shape of its API is public, while its model and training data are not.
That shape is simple. You send a block of text, called the state, plus questions of three types: a choice from a list, a score on an ordered scale, or a yes-or-no probability. Kev, Clef, ai_decide and Laya’s server all accept requests in that same format, so switching between many of these Jev alternatives can mean changing one web address.
Copying the accuracy is the hard part
Matching Jev’s answers is much harder for the open Jev alternatives. Laya’s own README is unusually honest about this: its base English checkpoint scores 0.362 on a typed-decisions benchmark, and only reaches 0.766 after fine-tuning on decisions from the same four workflows. Treat it as a fast starting point you train on your own examples.
Kev publishes the most careful comparison. On questions from sources it never trained on, its largest model, Kev-27B, scores 0.851 against Jev’s 0.857. Kev’s smaller models trail further, and on the harder MMLU-Pro knowledge test Kev-27B scores 0.675 against Jev’s 0.840. Kev’s author notes that because nobody knows what Jev was trained on, this “isn’t a controlled comparison”.
Decision Index score on held-out datasets, chance-corrected (Kev’s published table; bars scaled to Jev = 54.0)
How the Jev Alternatives Compare on Speed and Price
Speed and price are where Jev alternatives make their strongest claims. They are also where the numbers are hardest to compare, because each company measures its own model on its own hardware.
Speed depends on who is holding the stopwatch
Cloudflare’s launch post includes the most complete latency table. In its tests, Jev took a median of 524.1 milliseconds per decision, Clef took 209.3 and Clef-flash took 38.8. Laya was fastest at 5.8 milliseconds, though Cloudflare notes it “trades off quality”.
Median decision latency in Cloudflare’s own tests, milliseconds (bars scaled to Jev = 524.1)
These figures need care. Jev is only available as a hosted service, so its time almost certainly includes a trip across the internet, while some rivals were timed on local hardware. AutoTrust makes the same point in its own release, citing third-party measurements of 238 to 301 milliseconds for Jev’s hosted API and 137 milliseconds for JEV-27B running locally. Amazon reports about 115 milliseconds for Strands Decider on an Nvidia RTX 3090.
Price: per token, per GPU or per platform
The Jev alternatives fall into three pricing models. Hosted services charge per token: Jev at $0.042 per million input tokens with output free, and Clef at $0.24. OpenAI has not yet priced the Decisions API. Databricks bills ai_decide through its own platform.
Open-weight Jev alternatives cost nothing per call but need hardware. Kev’s author says a Kev-4B fine-tuning run costs about $1 on an H100, and Strands Decider runs on a laptop. For a team already paying for GPUs, the marginal cost of a decision can be close to zero.
The arithmetic also says something about TypeSafe. At list price, a trillion billed input tokens a day would be $42,000 a day, or about $15.3 million a year, by our calculation. The real figure could be lower after discounts and free use, which makes the reported $10 billion valuation a bet on growth rather than current revenue.
Benchmarks: Where Jev Still Beats the Jev Alternatives
Every one of the Jev alternatives in this race says it beats Jev on something. Read closely, the published tables show Jev still winning several important tests.
Cloudflare’s own table shows Jev winning some tests
Cloudflare’s benchmark selection is useful because it includes results that do not flatter Clef. The table shows six of the tests it published.
| Benchmark (what it tests) | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL, choosing the right tool call | 98.47 | 98.76 | 95.75 |
| BANKING77, sorting banking queries into 77 types | 94.20 | 90.93 | 79.74 |
| When2Call, knowing when to call a tool | 72.37 | 65.58 | 80.97 |
| BRIGHT, finding relevant documents | 45.91 | 39.26 | 47.52 |
| PhishNChips, spotting phishing | 79.60 | 75.05 | 62.55 |
| TypeSafe’s agent trace workflow | 68.5 | 69.8 | 71.6 |
Jev wins on When2Call, BRIGHT and TypeSafe’s own agent-trace workflow. Those are tasks that need judgement rather than sorting into fixed categories, which is where a model’s training matters most. Clef wins clearly on classification sets with many fixed categories.
Almost every benchmark for Jev alternatives is self-reported
The Register points out that Cloudflare’s scores have not yet been reproduced on the official Decision Index. AutoTrust and Fastino both ran their own comparisons. Kev’s comparison uses its own test suites. Until someone neutral runs all of these Jev alternatives on the same questions, the honest summary is that several are close to Jev on some tasks and none has clearly overtaken it.
Almeida says the copycats underestimate the work
Almeida told the Journal he was sure copycats would emerge and remains confident Jev will prove more capable. He was sharper with TechCrunch: “I get that people think it’s a gold rush, but they might be underestimating the difficulty of making the models actually smart.”
“The current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful,” he said. TypeSafe’s moat, he has argued before, is the synthetic training data it generates.
The older objection: is this just a classifier?
Not everyone accepts that decision models are new. Anastasios Angelopoulos, chief executive of the evaluation platform Arena, told the Financial Times: “It’s unclear to me what makes these models different from standard ‘zero-shot classifiers’, which are relatively well-known technology.” Meta, Google and Hugging Face have offered such tools for years. The defenders’ answer is calibration: a decision model’s probabilities are meant to be trustworthy enough to act on automatically.
Are Jev and the Jev Alternatives Really a Replacement for LLMs?
The Journal frames Jev as part of a debate about alternatives to large language models. Looking at how the Jev alternatives are built complicates that framing.
Most Jev alternatives are language models underneath
Amazon’s Strands Decider is a Qwen language model with its text-generating head removed. Cloudflare’s Clef runs a frozen Qwen model for one reading pass. AutoTrust’s JEV-27B adds a small block to a frozen Qwen and keeps the original’s ability to write code. Kev is a set of small adapters on Qwen. The Financial Times reported that TypeSafe itself trained Jev using open-weight models and synthetic data.
So the change is less about replacing language models than about using them differently. Instead of generating an answer word by word, these models read the input once and score a fixed list of answers. That is why they are fast, and why they cannot hallucinate an answer outside the list.
What Jev and the Jev alternatives cannot do
The builders of the Jev alternatives are candid about the limits. Amazon’s launch post says the approach makes its model “significantly worse at solving complex problems than reasoning models”, and unsuitable for coding, chat or summaries. AutoTrust warns that JEV-27B shares Jev’s weak spots: questions needing several reasoning steps, arithmetic, dates and deliberately tricky inputs.
Fastino’s GLiDE is a direct response to that weakness. It lets the model reason further on hard questions, and Fastino says about a third of requests in its tests needed that slower path. That makes it faster than a reasoning model on easy work but less predictable on timing.
The likely result is a hybrid
The pattern the builders describe is a mix. Amazon’s post says teams are “using LLMs to make the hardest decisions and using decider models to make the easier rote decisions, reducing cost and latency”. That is a useful split for any company building autonomous AI agents, where most steps are small routing choices.
Cloudflare goes further, writing that with decision models “a human does not necessarily need to be in the loop for agentic decisions anymore”. That claim deserves caution. A calibrated probability tells you how often the model is likely to be wrong; it does not remove the need to decide who is accountable when it is.
What the Jev Alternatives Mean for TypeSafe
The rapid arrival of Jev alternatives is a sign that the idea is good. It is also a warning about how long an early lead lasts.
Switching between Jev alternatives costs almost nothing
TypeSafe published a clean API, and the makers of the Jev alternatives adopted it. Kev works with TypeSafe’s own software development kit, Cloudflare calls Clef “fully Jev-API compatible”, and Databricks says ai_decide is “directly compatible”. A customer can test most Jev alternatives without rewriting code, which keeps pressure on TypeSafe’s price and quality.
Cloudflare even uses the same name for part of its training. Its post says it developed “Reinforcement Learning for Calibrated Decisions (RLCD)” as a secondary training target, the phrase the Journal says TypeSafe uses to describe its own method.
Where TypeSafe still differs
TypeSafe’s documentation confirms two differences. Jev accepts text only, with no image, audio or video input, and it is not fine-tuned on customer data: every account uses the same weights. It is also available only as a hosted API. Each of these is a gap a rival has chosen to fill, with image input from Clef and OpenAI, fine-tuning from Cloudflare and Kev, and local hosting from every open model.
For TypeSafe, the defence is accuracy and calibration. Its rivals’ own tables suggest it still leads on harder judgement tasks, at a lower price per token than Cloudflare. Whether that lead survives the next model releases from much larger companies is what investors are betting on.
How to Choose Between Jev Alternatives
For teams deciding whether to use Jev or one of the Jev alternatives, the right choice depends on where the data can go and how hard the decisions are. The table is a starting point, not a ranking.
| If you need | Options to test first | Main trade-off |
|---|---|---|
| Best accuracy on text, hosted | Jev, Clef | Data leaves your network |
| Decisions about images | Clef, OpenAI Decisions API | Higher price, or no published price |
| Data that cannot leave your servers | Kev-27B, JEV-27B, Clef open weights | You need a large data-centre GPU |
| Decisions inside a data warehouse | Databricks ai_decide | Tied to one platform, still in beta |
| Fast local experiments | Strands Decider 2B, Kev-4B, Laya | Lower accuracy without fine-tuning |
| Hard decisions with reasoning | Fastino GLiDE, or an LLM | Slower and less predictable timing |
Run a shadow test before switching
Do not trust any vendor’s benchmark, including TypeSafe’s. Send a copy of real traffic to the model you use now and to one or two Jev alternatives, log what each would have decided, and compare the answers with outcomes you can check. A week of shadow traffic will tell you more than any leaderboard. Our AI strategy team can help design that kind of test.
Check calibration, not just accuracy
The whole point of a decision model is acting on its confidence. Measure how often the model is wrong when it says it is 90% sure. Kev’s README reports that on questions with no knowable answer, Jev still answered with at least 0.9 confidence 9% of the time. Pick your automation threshold from your own data.
Pin the version you tested
Jev alternatives are changing weekly, and so is Jev. TypeSafe’s documentation warns that its “jev-latest” alias moves when a new model ships, so answers can change without any change on your side. Pin a specific version, log the model that produced each decision, and treat upgrades as a change that needs IT governance sign-off.
What to Watch Next for Jev Alternatives
Four things will show whether the Jev alternatives become a real market or fade.
Independent tests of the Jev alternatives
The official Decision Index needs to reproduce the scores that Cloudflare, Fastino and AutoTrust have reported for their Jev alternatives. If neutral tests confirm them, Jev will have serious competition on quality, not just price.
OpenAI’s price
OpenAI has not priced its Decisions API. If it charges close to Jev’s rate and adds image input, it will become the default for many companies already using OpenAI.
TypeSafe’s funding and next model
A round of $1 billion or more would give TypeSafe the money to train much larger models. Its next release will show whether it can stay ahead of rivals that copied its interface in a day.
Whether customers mix models
The most likely outcome is that companies use several Jev alternatives alongside Jev, choosing by task, as they already do with language models. A decision step that can be swapped without rewriting code is good for buyers, whoever wins.
Jev Alternatives FAQ
What is Jev?
Jev is a decision model from TypeSafe AI, released on 15 September 2026. It answers yes-or-no, multiple-choice and rating questions about text and returns a probability for each answer. It does not generate text.
What are the main Jev alternatives?
Hosted options include OpenAI’s Decisions API, Databricks’ ai_decide, Cloudflare’s Clef and Fastino’s GLiDE. Open-weight options include Amazon’s Strands Decider 2B, Cloudflare’s Clef weights, AutoTrust’s JEV-27B, Kev and Laya.
Are the Jev alternatives better than Jev?
Some claim wins on particular tests, but almost all results are self-reported. On published tables, Jev still leads on several harder judgement tasks, and the best open model, Kev-27B, scores slightly below it.
Are decision models a replacement for LLMs?
Not in general. They are faster and cheaper for choosing between fixed answers, but cannot write, code or reason through complex problems. Most are built on top of a language model.
How much does Jev cost?
TypeSafe charges $0.042 per million input tokens, and output is free. Cloudflare’s Clef costs $0.24 per million tokens.
References and Further Reading
Startup TypeSafe AI’s Jev Model Sparks Copycats, Talk of LLM Alternatives (The Wall Street Journal)
Amazon releases its own Jev clone as decision models flood the web (TechCrunch)
Introducing Clef: our open-source decision models (Cloudflare)
Cloudflare tries to outplay Jev with open-weight Clef models (The Register)
Introducing Strands Decider 2B (Strands Agents)
Small Decisions: Engineering a Leading Model (Marc Brooker)
Introducing ai_decide (Databricks)
AutoTrust AI Releases JEV-27B (PR Newswire)
Fastino Labs Releases GLiDE (PR Newswire via Yahoo Finance)
Jev Fervor Leads to Talk of Big Valuation Boost (The Information)
The cheap new AI model taking aim at OpenAI and Anthropic (Financial Times via Financial Post)
Kev: Jev-like decision models (GitHub)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.