Mistral Large 4 ranking on the Artificial Analysis Intelligence Index is sixth among open-weight models on the chart Artificial Analysis published at launch, and eighth on its full leaderboard. Both numbers are right. They count different things, and the gap between them says a lot about how to read AI leaderboards before you choose a model.
The model scored 38.4 on the index when Artificial Analysis tested the preview on 6 October 2026, the highest score for any model built outside the US and China. We covered the model itself, its specifications, price and self-hosting needs, in our Mistral Large 4 launch report. This article is about the scoreboard.
Below we show exactly how the Mistral Large 4 ranking is counted, why the preview is not yet officially listed as open, what the index measures, where Mistral’s model is strong and weak against its open rivals, what each point costs, and what could move it before the weights ship.
Table of contents
- Sixth or Eighth? How the Mistral Large 4 Ranking Is Counted
- Why the Preview Is Not Yet Officially Open
- What the Intelligence Index Measures
- How Far Behind, in Points
- Cost and Speed per Task
- The Cyber Index and Other Leaderboards
- How to Read a Mistral Large 4 Ranking Before You Buy
- What Could Move the Mistral Large 4 Ranking by November
- Mistral Large 4 Ranking: Frequently Asked Questions
- References
Sixth or Eighth? How the Mistral Large 4 Ranking Is Counted
Mistral Large 4 sits at 38.4 on a scale that runs to 100. Where that lands it depends on which models you put on the list.
The launch chart: five open models ahead
Artificial Analysis announced the result with a chart of 24 leading models. Five open-weight models on that chart score higher: Xiaomi’s MiMo-V2.6-Pro, Z.ai’s GLM-5.3 and GLM-5.3-Flash, Moonshot AI’s Kimi K3 and DeepSeek V4.1 Flash. That makes the Mistral Large 4 ranking sixth among the open models shown, which is where the “6th place” headline comes from.
The full leaderboard: seven ahead
The chart is a selection, not the whole table. Artificial Analysis’s complete data also includes two Alibaba open models that score above Mistral: Qwen3.8 2.4T A95B at 39.9 and Qwen3.8-Flash-Next at 39.8. Add them and the Mistral Large 4 ranking becomes eighth, the figure Trending Topics reported, with all seven models ahead of it from Chinese labs.
| Place | Open-weight model | Lab | Index score | On launch chart? |
|---|---|---|---|---|
| 1 | MiMo-V2.6-Pro | Xiaomi | 46.3 | Yes |
| 2 | GLM-5.3 (Max) | Z.ai | 44.8 | Yes |
| 3 | Kimi K3 (Max) | Moonshot AI | 43.6 | Yes |
| 4 | GLM-5.3-Flash | Z.ai | 41.8 | Yes |
| 5 | Qwen3.8 2.4T A95B | Alibaba | 39.9 | No |
| 6 | Qwen3.8-Flash-Next | Alibaba | 39.8 | No |
| 7 | DeepSeek V4.1 Flash (Max) | DeepSeek | 39.5 | Yes |
| 8 | Mistral Large 4 Preview | Mistral | 38.4 | Yes |
| 9 | MiMo-V2.6-Flash | Xiaomi | 37.9 | No |
| 10 | DeepSeek V4 Pro 0813 (Max) | DeepSeek | 36.0 | No |
Scores are from Artificial Analysis’s model data on 7 October, rounded to one decimal place. The chart on 6 October showed the same models as rounded whole numbers.
Why the launch chart showed fewer models
Artificial Analysis does not explain how it picks the models on a launch chart, but the selection is visible. Alibaba appears twice: as Qwen3.8 Max (0902), a proprietary model that scored 45, and as the open Qwen3.8 27B, which scored 34. Its two larger open models, which sit just above Mistral, were simply not on the chart. That one editorial choice is the whole difference between a sixth-place and an eighth-place Mistral Large 4 ranking.
Counting labs instead of models
There is a third way to count, and it also gives sixth. Z.ai and Alibaba each have two models above Mistral, so only five labs are ahead: Xiaomi, Z.ai, Moonshot AI, Alibaba and DeepSeek. Ranked by each lab’s best open model, the Mistral Large 4 ranking is sixth, and Mistral is the only non-Chinese lab in the top six.
Every way of counting, side by side
| Way of counting | Open entries ahead | Mistral Large 4 position |
|---|---|---|
| Open models on the launch chart | 5 | 6th |
| Full open-weight leaderboard, by model | 7 | 8th |
| Best model per lab | 5 | 6th |
| Permissively licensed open models only, if Mistral’s licence is permissive | 3 | 4th |
| Open models from outside China | 0 | 1st |
The licence row uses Artificial Analysis’s own labels: of the seven models ahead, MiMo-V2.6-Pro, GLM-5.3-Flash and DeepSeek V4.1 Flash carry the MIT licence, and the other four use custom licences that Artificial Analysis files under commercial-use restrictions.
The other sixth: legal work
There is one more sixth place worth knowing. On Harvey’s Legal Agent Benchmark, run by evaluator Vals AI, Mistral Large 4 placed sixth of 75 models with 15.83%, the best result for any open-weight model. Vals AI also ranked it ninth among open models on its broader Vals Index, so the strong legal score is a speciality, not the general picture.
Why the Preview Is Not Yet Officially Open
Strictly speaking, every open-weight Mistral Large 4 ranking above is provisional. Artificial Analysis currently labels the model a “Proprietary model”.
What the label means today
Mistral released Mistral Large 4 as a “Research Public Preview” on its API, with weights promised “by the end of the month”. Until those weights are downloadable, Artificial Analysis lists no licence and no weights link, so the preview cannot appear when its site is filtered to open-weight models. As Trending Topics noted, the model joins the open ranking only once the weights are released.
What changes on release day
When the weights ship, three things are set at once: the licence, the weights category and, possibly, a fresh score. If reinforcement learning has continued, as Mistral says it will, Artificial Analysis may re-test the final model, and the Mistral Large 4 ranking could move either way.
Why the licence label matters
Artificial Analysis sorts open models into “Open Weights”, “Open Weights (Commercial Use Restricted)” and “Open Weights (Non-commercial)”. VentureBeat has reported a custom Mistral licence rather than Apache 2.0. If that licence restricts commercial use, buyers who filter for permissive terms will see a different Mistral Large 4 ranking from those who do not.
What the Intelligence Index Measures
A single score hides ten tests. Version 4.3.2 of the Artificial Analysis Intelligence Index combines AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR, covering knowledge work, agents, coding, science and long documents. Every Mistral Large 4 ranking in this article rests on those ten results.
| Evaluation | Mistral Large 4 | MiMo-V2.6-Pro | GLM-5.3 (Max) | DeepSeek V4.1 Flash |
|---|---|---|---|---|
| Humanity’s Last Exam | 35.0% | 49.4% | 42.3% | 39.2% |
| SciCode | 54.2% | 60.9% | 59.0% | 51.9% |
| AA-LCR, long documents | 81.3% | 86.3% | 79.7% | 84.0% |
| Terminal-Bench 4.0 | 26.8% | 34.8% | 41.9% | 26.8% |
| CritPt, research physics | 10.6% | 26.6% | 19.1% | 14.3% |
| GDP.pdf, documents and images | 18.6% | 19.2% | 11.2% | 12.8% |
| AutomationBench-AA | 59.9% | 58.6% | 62.2% | 68.9% |
| GDPval-AA, Elo rating | 1,424 | 1,686 | 1,653 | 1,600 |
| Hallucination rate, AA-Omniscience | 41.9% | 40.6% | 29.6% | 96.5% |
All figures are from Artificial Analysis’s model data. Lower is better only for the hallucination rate.
Where Mistral Large 4 is strong
The preview does best on documents. Its 18.6% on GDP.pdf is level with MiMo-V2.6-Pro and well ahead of GLM-5.3, which Artificial Analysis linked partly to Mistral’s API now accepting 100 images per request, up from eight. On long-document reading it beats GLM-5.3, and on AutomationBench it edges past MiMo-V2.6-Pro by 1.3 points.
Where it trails
The weak spots are hard reasoning. On Humanity’s Last Exam it is the lowest of the four, 14.4 points behind MiMo-V2.6-Pro (49.4 minus 35.0), and on CritPt it scores 10.6% against MiMo’s 26.6%. On GDPval-AA, a test of real professional tasks, it sits 262 Elo points behind the leader (1,686 minus 1,424). These gaps decide the Mistral Large 4 ranking more than any single strength.
Fewer confident mistakes than some rivals
The hallucination column matters for business use. Mistral Large 4 makes up an answer, rather than declining, 41.9% of the time it does not know, close to MiMo-V2.6-Pro. DeepSeek V4.1 Flash, which sits one place higher, does so 96.5% of the time. A higher place in a Mistral Large 4 ranking table does not always mean the safer model.
How Far Behind, in Points
Places hide distances, and the Mistral Large 4 ranking is a good example. On this index, several models are packed within two points of each other, so a small improvement can move Mistral Large 4 up several places at once.
Bar widths equal the score on the index’s 0 to 100 scale.
What it would take to climb
Using the unrounded scores, Mistral Large 4 is 1.08 points behind DeepSeek V4.1 Flash (39.46 minus 38.38), 1.51 points behind Qwen3.8 2.4T, and 3.43 points behind GLM-5.3-Flash. So a gain of about 1.6 points would lift the Mistral Large 4 ranking from eighth to fifth on the full table, while fourth needs nearly three and a half.
The lead over the rest of the world
Behind it, the gap is wider. Before this release, the strongest open model from outside China was South Korea’s Motif 3 at 33.6, so Mistral has moved the non-Chinese best up by 4.8 points in a single release. Its own predecessor, Mistral Large 3, scored 9.3, a jump of about 29 points.
Cost and Speed per Task
A Mistral Large 4 ranking says nothing about the bill. Artificial Analysis also records what each model cost to run the whole index, divided into a cost per task, and how long each task took.
| Model | Cost per task | Seconds per task | Index points per dollar |
|---|---|---|---|
| MiMo-V2.6-Pro | $0.13 | 1,271 | 348 |
| GLM-5.3-Flash | $0.25 | 979 | 165 |
| DeepSeek V4.1 Flash (Max) | $0.27 | 285 | 149 |
| Qwen3.8-Flash-Next | $0.37 | 1,444 | 107 |
| Mistral Large 4, launch discount | $0.57 | 507 | 67 |
| Mistral Large 4, list price | $1.13 | 507 | 34 |
| Kimi K3 (Max) | $2.00 | 966 | 22 |
| GLM-5.3 (Max) | $2.01 | 955 | 22 |
| Qwen3.8 2.4T A95B | $2.16 | 1,425 | 19 |
Points per dollar is the index score divided by the cost per task, for example 38.38 divided by 1.13 is about 34. Costs use each provider’s standard API price.
Over four times the flash models
Artificial Analysis put it plainly: Mistral Large 4 has “over 4x the Cost per Task of similar-intelligence open weights models”. At $1.13 it costs about 4.2 times DeepSeek V4.1 Flash ($1.13 divided by $0.27) for a slightly lower score. The two-week launch discount halves that to $0.57, still more than twice the flash models.
Cheaper than the big Chinese models
The comparison flips against the largest rivals. GLM-5.3, Kimi K3 and Qwen3.8 2.4T all cost about $2 a task, so Mistral Large 4 is roughly 44% cheaper than GLM-5.3 (1.13 divided by 2.01 is 0.56) while scoring 6.4 points lower. Whether that trade suits you depends on the task, which is why a Mistral Large 4 ranking alone should never settle a buying decision.
Second fastest of the top eight
Speed is a genuine strength. At about 507 seconds per index task, Mistral Large 4 is the second fastest of the eight leading open models, behind only DeepSeek V4.1 Flash at 285 seconds and well ahead of MiMo-V2.6-Pro at 1,271. For interactive tools, that can matter more than a point or two on the index, and it is a strength no Mistral Large 4 ranking table shows.
The Cyber Index and Other Leaderboards
The Intelligence Index is not the only scoreboard, and the Mistral Large 4 ranking looks different on each. Mistral built its launch around a different one.
Top three on cyber, once open
Mistral Large 4 scored 50 on the Artificial Analysis Cyber Index, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56. Artificial Analysis says that “once its weights are released, it will rank among the top three open weights models on the Cyber Index”. Its best single result is 82% on CyberGym-E2E-AA, ahead of MiMo-V2.6-Pro at 79%.
Vals AI’s view
| Leaderboard | Score | Rank |
|---|---|---|
| Artificial Analysis Intelligence Index | 38.4 | 6th on launch chart, 8th on full open list |
| Artificial Analysis Cyber Index | 50 | Top three open, once weights ship |
| Harvey’s Legal Agent Benchmark (Vals AI) | 15.83% | 6th of 75, 1st open |
| Vals Index | 48.05% | 32nd of 44, 9th open |
| Finance Agent v2 (Vals AI) | 54.68% | 22nd of 75 |
| Terminal-Bench 4.0 (Vals AI) | 22.73% | 20th of 44 |
The pattern is consistent: a specialist strong in legal, document and cyber work, and middling on general reasoning and coding. Each Mistral Large 4 ranking depends on which of those the leaderboard weights most.
How to Read a Mistral Large 4 Ranking Before You Buy
Leaderboards are useful for shortlisting and poor for final choices. Five checks turn a Mistral Large 4 ranking into a decision.
Match the benchmark to the job
If your work is contracts, reports and scanned documents, the legal and GDP.pdf results matter more than the overall Mistral Large 4 ranking. If it is research-grade maths or agentic coding, the Humanity’s Last Exam and Terminal-Bench gaps matter more. Pick the two or three tests closest to your workload and compare models on those.
Watch the size of the model
Size matters for anyone planning to self-host. Mistral Large 4 has 1 trillion total parameters with 49 billion active, yet three far smaller open models score higher: Qwen3.8-Flash-Next (180 billion total, 6 billion active), GLM-5.3-Flash (320 billion, 18 billion) and DeepSeek V4.1 Flash (552 billion, 16 billion). On points per gigabyte of memory, the Mistral Large 4 ranking would be much lower than its index place suggests.
Price the finished task, not the token
Cost per task already reflects how many tokens a model writes. Mistral Large 4 is verbose, so its per-token price looks better than its per-task cost. Our LLM API pricing guide walks through the method, and our guide to open-weight AI models in 2026 covers the wider field.
Check the licence before the leaderboard
For an open model, the licence decides what you may build. A model can top a Mistral Large 4 ranking table and still be unusable for your product if its terms restrict commercial use. Read the licence text on release day, before you plan around the weights.
Run your own test
Run 50 to 100 of your own tasks through Mistral Large 4 and one or two rivals, and record accuracy, cost and time. Our AI strategy team builds exactly these pilot evaluations for UK firms choosing between open and closed models.
What Could Move the Mistral Large 4 Ranking by November
The current Mistral Large 4 ranking is a snapshot of a model still in training. Three things could change its place within weeks.
The final model and its weights
Mistral says the model “continues to improve rapidly as we refine it” because reinforcement learning is still running. A re-test of the released model could close the 1.08-point gap to seventh, or more. The weights, due by the end of October, also move Mistral Large 4 from the proprietary list onto the open one for good.
New releases from rivals
The open-weight field moves fast. Reflection’s Beam is due to release weights later in October, and Chinese labs have shipped new versions every few weeks this year. Each new entry above 38.4 pushes the Mistral Large 4 ranking down a place, whatever Mistral does.
Changes to the index itself
Artificial Analysis updates its index as tests saturate; the current version is 4.3.2. A future version that weights document or cyber work more heavily would favour Mistral, while more hard-science tests would not. Compare scores only within the same index version.
Mistral Large 4 Ranking: Frequently Asked Questions
What place is Mistral Large 4 on the Artificial Analysis index?
It scored 38.4. That is sixth among the open-weight models on Artificial Analysis’s launch chart, eighth on its full open-weight leaderboard, and first among open models from outside China.
Why do some reports say sixth and others eighth?
The launch chart left out two Alibaba Qwen3.8 models that score slightly higher. Counting by lab rather than by model also gives sixth, because two labs have two models each above Mistral.
Is Mistral Large 4 officially an open-weight model yet?
Not yet. Artificial Analysis lists the preview as proprietary until Mistral releases the weights, which it has promised by the end of October 2026.
Which open model is ranked first?
Xiaomi’s MiMo-V2.6-Pro, at 46.3 on the Artificial Analysis Intelligence Index, 7.9 points above Mistral Large 4.
Does a higher ranking mean a better model for my business?
Not necessarily. A Mistral Large 4 ranking blends ten tests into one number. If your work matches its strengths, such as legal documents, long reports and cyber defence, it can beat higher-ranked models on the tasks you care about.
Is Mistral Large 4 good value?
It costs $1.13 per index task at list price, about four times the cheapest models of similar score, but roughly 44% less than GLM-5.3 and Kimi K3. It is also one of the fastest.
References
Mistral Large 4 Preview: intelligence, performance and price analysis (Artificial Analysis)
Mistral Large 4 scores 38 on the Intelligence Index (Artificial Analysis on X)
Mistral Large 4 trails China’s open-weight leaders on Artificial Analysis (Trending Topics)
Mistral Large 4 becomes top non-Chinese open model on Artificial Analysis (OfficeChai)
Mistral Large 4 benchmarks, cost and capabilities (Vals AI)
Mistral Large 4 is the top open-weight model on HLAB (Vals AI on X)
Mistral debuts Large 4 Le Chonk (VentureBeat)
Mistral launches Large 4 preview with 1T parameters (TestingCatalog)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.