Mistral Large 4 ranking on the Artificial Analysis Intelligence Index is sixth among open-weight models on the chart Artificial Analysis published at launch, and eighth on its full leaderboard. Both numbers are right. They count different things, and the gap between them says a lot about how to read AI leaderboards before you choose a model.

The model scored 38.4 on the index when Artificial Analysis tested the preview on 6 October 2026, the highest score for any model built outside the US and China. We covered the model itself, its specifications, price and self-hosting needs, in our Mistral Large 4 launch report. This article is about the scoreboard.

Below we show exactly how the Mistral Large 4 ranking is counted, why the preview is not yet officially listed as open, what the index measures, where Mistral’s model is strong and weak against its open rivals, what each point costs, and what could move it before the weights ship.

Sixth or Eighth? How the Mistral Large 4 Ranking Is Counted

mistral large 4 ranking artificial analysis open weight index b stadium grandstand with tiered rows of seats

Mistral Large 4 sits at 38.4 on a scale that runs to 100. Where that lands it depends on which models you put on the list.

The launch chart: five open models ahead

Artificial Analysis announced the result with a chart of 24 leading models. Five open-weight models on that chart score higher: Xiaomi’s MiMo-V2.6-Pro, Z.ai’s GLM-5.3 and GLM-5.3-Flash, Moonshot AI’s Kimi K3 and DeepSeek V4.1 Flash. That makes the Mistral Large 4 ranking sixth among the open models shown, which is where the “6th place” headline comes from.

The full leaderboard: seven ahead

The chart is a selection, not the whole table. Artificial Analysis’s complete data also includes two Alibaba open models that score above Mistral: Qwen3.8 2.4T A95B at 39.9 and Qwen3.8-Flash-Next at 39.8. Add them and the Mistral Large 4 ranking becomes eighth, the figure Trending Topics reported, with all seven models ahead of it from Chinese labs.

PlaceOpen-weight modelLabIndex scoreOn launch chart?
1MiMo-V2.6-ProXiaomi46.3Yes
2GLM-5.3 (Max)Z.ai44.8Yes
3Kimi K3 (Max)Moonshot AI43.6Yes
4GLM-5.3-FlashZ.ai41.8Yes
5Qwen3.8 2.4T A95BAlibaba39.9No
6Qwen3.8-Flash-NextAlibaba39.8No
7DeepSeek V4.1 Flash (Max)DeepSeek39.5Yes
8Mistral Large 4 PreviewMistral38.4Yes
9MiMo-V2.6-FlashXiaomi37.9No
10DeepSeek V4 Pro 0813 (Max)DeepSeek36.0No

Scores are from Artificial Analysis’s model data on 7 October, rounded to one decimal place. The chart on 6 October showed the same models as rounded whole numbers.

Why the launch chart showed fewer models

Artificial Analysis does not explain how it picks the models on a launch chart, but the selection is visible. Alibaba appears twice: as Qwen3.8 Max (0902), a proprietary model that scored 45, and as the open Qwen3.8 27B, which scored 34. Its two larger open models, which sit just above Mistral, were simply not on the chart. That one editorial choice is the whole difference between a sixth-place and an eighth-place Mistral Large 4 ranking.

Counting labs instead of models

There is a third way to count, and it also gives sixth. Z.ai and Alibaba each have two models above Mistral, so only five labs are ahead: Xiaomi, Z.ai, Moonshot AI, Alibaba and DeepSeek. Ranked by each lab’s best open model, the Mistral Large 4 ranking is sixth, and Mistral is the only non-Chinese lab in the top six.

Every way of counting, side by side

Way of countingOpen entries aheadMistral Large 4 position
Open models on the launch chart56th
Full open-weight leaderboard, by model78th
Best model per lab56th
Permissively licensed open models only, if Mistral’s licence is permissive34th
Open models from outside China01st

The licence row uses Artificial Analysis’s own labels: of the seven models ahead, MiMo-V2.6-Pro, GLM-5.3-Flash and DeepSeek V4.1 Flash carry the MIT licence, and the other four use custom licences that Artificial Analysis files under commercial-use restrictions.

The other sixth: legal work

There is one more sixth place worth knowing. On Harvey’s Legal Agent Benchmark, run by evaluator Vals AI, Mistral Large 4 placed sixth of 75 models with 15.83%, the best result for any open-weight model. Vals AI also ranked it ninth among open models on its broader Vals Index, so the strong legal score is a speciality, not the general picture.

Why the Preview Is Not Yet Officially Open

mistral large 4 ranking artificial analysis open weight index c bobsleigh in a banked ice track

Strictly speaking, every open-weight Mistral Large 4 ranking above is provisional. Artificial Analysis currently labels the model a “Proprietary model”.

What the label means today

Mistral released Mistral Large 4 as a “Research Public Preview” on its API, with weights promised “by the end of the month”. Until those weights are downloadable, Artificial Analysis lists no licence and no weights link, so the preview cannot appear when its site is filtered to open-weight models. As Trending Topics noted, the model joins the open ranking only once the weights are released.

What changes on release day

When the weights ship, three things are set at once: the licence, the weights category and, possibly, a fresh score. If reinforcement learning has continued, as Mistral says it will, Artificial Analysis may re-test the final model, and the Mistral Large 4 ranking could move either way.

Why the licence label matters

Artificial Analysis sorts open models into “Open Weights”, “Open Weights (Commercial Use Restricted)” and “Open Weights (Non-commercial)”. VentureBeat has reported a custom Mistral licence rather than Apache 2.0. If that licence restricts commercial use, buyers who filter for permissive terms will see a different Mistral Large 4 ranking from those who do not.

What the Intelligence Index Measures

mistral large 4 ranking artificial analysis open weight index d shot put circle with distance marker flags

A single score hides ten tests. Version 4.3.2 of the Artificial Analysis Intelligence Index combines AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR, covering knowledge work, agents, coding, science and long documents. Every Mistral Large 4 ranking in this article rests on those ten results.

EvaluationMistral Large 4MiMo-V2.6-ProGLM-5.3 (Max)DeepSeek V4.1 Flash
Humanity’s Last Exam35.0%49.4%42.3%39.2%
SciCode54.2%60.9%59.0%51.9%
AA-LCR, long documents81.3%86.3%79.7%84.0%
Terminal-Bench 4.026.8%34.8%41.9%26.8%
CritPt, research physics10.6%26.6%19.1%14.3%
GDP.pdf, documents and images18.6%19.2%11.2%12.8%
AutomationBench-AA59.9%58.6%62.2%68.9%
GDPval-AA, Elo rating1,4241,6861,6531,600
Hallucination rate, AA-Omniscience41.9%40.6%29.6%96.5%

All figures are from Artificial Analysis’s model data. Lower is better only for the hallucination rate.

Where Mistral Large 4 is strong

The preview does best on documents. Its 18.6% on GDP.pdf is level with MiMo-V2.6-Pro and well ahead of GLM-5.3, which Artificial Analysis linked partly to Mistral’s API now accepting 100 images per request, up from eight. On long-document reading it beats GLM-5.3, and on AutomationBench it edges past MiMo-V2.6-Pro by 1.3 points.

Where it trails

The weak spots are hard reasoning. On Humanity’s Last Exam it is the lowest of the four, 14.4 points behind MiMo-V2.6-Pro (49.4 minus 35.0), and on CritPt it scores 10.6% against MiMo’s 26.6%. On GDPval-AA, a test of real professional tasks, it sits 262 Elo points behind the leader (1,686 minus 1,424). These gaps decide the Mistral Large 4 ranking more than any single strength.

Fewer confident mistakes than some rivals

The hallucination column matters for business use. Mistral Large 4 makes up an answer, rather than declining, 41.9% of the time it does not know, close to MiMo-V2.6-Pro. DeepSeek V4.1 Flash, which sits one place higher, does so 96.5% of the time. A higher place in a Mistral Large 4 ranking table does not always mean the safer model.

How Far Behind, in Points

mistral large 4 ranking artificial analysis open weight index e ski jump ramp with a landing hill

Places hide distances, and the Mistral Large 4 ranking is a good example. On this index, several models are packed within two points of each other, so a small improvement can move Mistral Large 4 up several places at once.

Artificial Analysis Intelligence Index, open-weight models (0 to 100)
MiMo-V2.6-Pro 46.3
GLM-5.3 (Max) 44.8
Kimi K3 (Max) 43.6
GLM-5.3-Flash 41.8
Qwen3.8 2.4T A95B 39.9
Qwen3.8-Flash-Next 39.8
DeepSeek V4.1 Flash (Max) 39.5
Mistral Large 4 Preview 38.4
MiMo-V2.6-Flash 37.9
Motif 3 (South Korea) 33.6

Bar widths equal the score on the index’s 0 to 100 scale.

What it would take to climb

Using the unrounded scores, Mistral Large 4 is 1.08 points behind DeepSeek V4.1 Flash (39.46 minus 38.38), 1.51 points behind Qwen3.8 2.4T, and 3.43 points behind GLM-5.3-Flash. So a gain of about 1.6 points would lift the Mistral Large 4 ranking from eighth to fifth on the full table, while fourth needs nearly three and a half.

The lead over the rest of the world

Behind it, the gap is wider. Before this release, the strongest open model from outside China was South Korea’s Motif 3 at 33.6, so Mistral has moved the non-Chinese best up by 4.8 points in a single release. Its own predecessor, Mistral Large 3, scored 9.3, a jump of about 29 points.

Cost and Speed per Task

mistral large 4 ranking artificial analysis open weight index f hanging spring scale with a hook

A Mistral Large 4 ranking says nothing about the bill. Artificial Analysis also records what each model cost to run the whole index, divided into a cost per task, and how long each task took.

ModelCost per taskSeconds per taskIndex points per dollar
MiMo-V2.6-Pro$0.131,271348
GLM-5.3-Flash$0.25979165
DeepSeek V4.1 Flash (Max)$0.27285149
Qwen3.8-Flash-Next$0.371,444107
Mistral Large 4, launch discount$0.5750767
Mistral Large 4, list price$1.1350734
Kimi K3 (Max)$2.0096622
GLM-5.3 (Max)$2.0195522
Qwen3.8 2.4T A95B$2.161,42519

Points per dollar is the index score divided by the cost per task, for example 38.38 divided by 1.13 is about 34. Costs use each provider’s standard API price.

Over four times the flash models

Artificial Analysis put it plainly: Mistral Large 4 has “over 4x the Cost per Task of similar-intelligence open weights models”. At $1.13 it costs about 4.2 times DeepSeek V4.1 Flash ($1.13 divided by $0.27) for a slightly lower score. The two-week launch discount halves that to $0.57, still more than twice the flash models.

Cheaper than the big Chinese models

The comparison flips against the largest rivals. GLM-5.3, Kimi K3 and Qwen3.8 2.4T all cost about $2 a task, so Mistral Large 4 is roughly 44% cheaper than GLM-5.3 (1.13 divided by 2.01 is 0.56) while scoring 6.4 points lower. Whether that trade suits you depends on the task, which is why a Mistral Large 4 ranking alone should never settle a buying decision.

Second fastest of the top eight

Speed is a genuine strength. At about 507 seconds per index task, Mistral Large 4 is the second fastest of the eight leading open models, behind only DeepSeek V4.1 Flash at 285 seconds and well ahead of MiMo-V2.6-Pro at 1,271. For interactive tools, that can matter more than a point or two on the index, and it is a strength no Mistral Large 4 ranking table shows.

The Cyber Index and Other Leaderboards

The Intelligence Index is not the only scoreboard, and the Mistral Large 4 ranking looks different on each. Mistral built its launch around a different one.

Top three on cyber, once open

Mistral Large 4 scored 50 on the Artificial Analysis Cyber Index, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56. Artificial Analysis says that “once its weights are released, it will rank among the top three open weights models on the Cyber Index”. Its best single result is 82% on CyberGym-E2E-AA, ahead of MiMo-V2.6-Pro at 79%.

Vals AI’s view

LeaderboardScoreRank
Artificial Analysis Intelligence Index38.46th on launch chart, 8th on full open list
Artificial Analysis Cyber Index50Top three open, once weights ship
Harvey’s Legal Agent Benchmark (Vals AI)15.83%6th of 75, 1st open
Vals Index48.05%32nd of 44, 9th open
Finance Agent v2 (Vals AI)54.68%22nd of 75
Terminal-Bench 4.0 (Vals AI)22.73%20th of 44

The pattern is consistent: a specialist strong in legal, document and cyber work, and middling on general reasoning and coding. Each Mistral Large 4 ranking depends on which of those the leaderboard weights most.

How to Read a Mistral Large 4 Ranking Before You Buy

Leaderboards are useful for shortlisting and poor for final choices. Five checks turn a Mistral Large 4 ranking into a decision.

Match the benchmark to the job

If your work is contracts, reports and scanned documents, the legal and GDP.pdf results matter more than the overall Mistral Large 4 ranking. If it is research-grade maths or agentic coding, the Humanity’s Last Exam and Terminal-Bench gaps matter more. Pick the two or three tests closest to your workload and compare models on those.

Watch the size of the model

Size matters for anyone planning to self-host. Mistral Large 4 has 1 trillion total parameters with 49 billion active, yet three far smaller open models score higher: Qwen3.8-Flash-Next (180 billion total, 6 billion active), GLM-5.3-Flash (320 billion, 18 billion) and DeepSeek V4.1 Flash (552 billion, 16 billion). On points per gigabyte of memory, the Mistral Large 4 ranking would be much lower than its index place suggests.

Price the finished task, not the token

Cost per task already reflects how many tokens a model writes. Mistral Large 4 is verbose, so its per-token price looks better than its per-task cost. Our LLM API pricing guide walks through the method, and our guide to open-weight AI models in 2026 covers the wider field.

Check the licence before the leaderboard

For an open model, the licence decides what you may build. A model can top a Mistral Large 4 ranking table and still be unusable for your product if its terms restrict commercial use. Read the licence text on release day, before you plan around the weights.

Run your own test

Run 50 to 100 of your own tasks through Mistral Large 4 and one or two rivals, and record accuracy, cost and time. Our AI strategy team builds exactly these pilot evaluations for UK firms choosing between open and closed models.

What Could Move the Mistral Large 4 Ranking by November

The current Mistral Large 4 ranking is a snapshot of a model still in training. Three things could change its place within weeks.

The final model and its weights

Mistral says the model “continues to improve rapidly as we refine it” because reinforcement learning is still running. A re-test of the released model could close the 1.08-point gap to seventh, or more. The weights, due by the end of October, also move Mistral Large 4 from the proprietary list onto the open one for good.

New releases from rivals

The open-weight field moves fast. Reflection’s Beam is due to release weights later in October, and Chinese labs have shipped new versions every few weeks this year. Each new entry above 38.4 pushes the Mistral Large 4 ranking down a place, whatever Mistral does.

Changes to the index itself

Artificial Analysis updates its index as tests saturate; the current version is 4.3.2. A future version that weights document or cyber work more heavily would favour Mistral, while more hard-science tests would not. Compare scores only within the same index version.

Mistral Large 4 Ranking: Frequently Asked Questions

What place is Mistral Large 4 on the Artificial Analysis index?

It scored 38.4. That is sixth among the open-weight models on Artificial Analysis’s launch chart, eighth on its full open-weight leaderboard, and first among open models from outside China.

Why do some reports say sixth and others eighth?

The launch chart left out two Alibaba Qwen3.8 models that score slightly higher. Counting by lab rather than by model also gives sixth, because two labs have two models each above Mistral.

Is Mistral Large 4 officially an open-weight model yet?

Not yet. Artificial Analysis lists the preview as proprietary until Mistral releases the weights, which it has promised by the end of October 2026.

Which open model is ranked first?

Xiaomi’s MiMo-V2.6-Pro, at 46.3 on the Artificial Analysis Intelligence Index, 7.9 points above Mistral Large 4.

Does a higher ranking mean a better model for my business?

Not necessarily. A Mistral Large 4 ranking blends ten tests into one number. If your work matches its strengths, such as legal documents, long reports and cyber defence, it can beat higher-ranked models on the tasks you care about.

Is Mistral Large 4 good value?

It costs $1.13 per index task at list price, about four times the cheapest models of similar score, but roughly 44% less than GLM-5.3 and Kimi K3. It is also one of the fastest.

References