Open-weight AI models have finally caught the frontier. In July 2026 Moonshot AI announced Kimi K3, and within weeks a model anyone can download was sitting at #3 on the Artificial Analysis Intelligence Index with a score of 57 — comparable to Claude Opus 4.8 and GPT-5.5, behind only Claude Fable 5 and GPT-5.6 Sol. For the first time in the short history of modern artificial intelligence, downloadable weights rank beside the best closed systems in the world.
Yet “open” has never meant so many different things at once. The label now stretches from genuine MIT and Apache 2.0 releases to revenue-gated bespoke licences, and from a 1.56TB datacentre download to a 30-billion-parameter agent that lives happily on a 24GB consumer graphics card. Treating every release as interchangeable is the fastest way to pick a model your legal team will later veto or your hardware simply cannot host.
That is why this guide ranks the eleven confirmed open-weight AI models of August 2026 by what you can actually run and legally ship — a licence tier times hardware tier decision matrix — rather than by leaderboard position alone. It sits alongside the deep dives in our AI models and tools hub, which covers each release in far more detail than a roundup can.
If your organisation is deciding whether self-hosted machine learning belongs in its stack at all, our AI strategy consultants work through exactly these licence, cost and infrastructure trade-offs every week. Consider this article the map; the territory changes monthly, and every figure below is dated and sourced so you can check it has not moved.
Table of contents
- Why Open-Weight AI Models Matter More Than Ever in 2026
- How to Judge Open-Weight AI Models: Licence Tiers Meet Hardware Tiers
- The Eleven Best Open-Weight AI Models at a Glance
- Frontier-Class Open-Weight AI Models: Kimi K3 and Qwen3.8-Max
- Production-Grade Open-Weight AI Models: GLM-5.2, DeepSeek V4 and MiniMax M3
- American Open-Weight AI Models: Inkling, Muse Glimmer and gpt-oss
- Small, Fast and European: The Rest of the Field
- Benchmarks: How the Leading Open Weights Stack Up
- API Pricing and Running Costs
- Licensing Deep Dive: Reading the Small Print
- Hardware: What It Takes to Run Them Locally
- How to Choose the Right Model for Your Team
- The Road Ahead for Open Weights
- Frequently Asked Questions
- References
Why Open-Weight AI Models Matter More Than Ever in 2026
The economics of frontier AI shifted decisively this summer. Open-weight AI models used to trail the closed frontier by a year or more; today the gap is measured in single-digit index points. That changes procurement conversations: when downloadable weights score within touching distance of the best proprietary systems, paying a premium for a closed API needs a justification beyond raw capability.
Control is the second driver. Open-weight AI models can be fine-tuned on private data, pinned to a version forever, audited for behaviour and run inside your own security perimeter. None of that is possible when a vendor can retire or retrain your model overnight. For regulated industries and sovereignty-minded governments, that difference is often decisive on its own.
The Month Open-Weight AI Models Caught the Frontier
Two independent measurements anchor the claim. Artificial Analysis places Kimi K3 at 57 on its Intelligence Index — third overall, comparable to Claude Opus 4.8 and GPT-5.5. Separately, the LLM Stats Open LLM Leaderboard of 7 August 2026 makes Kimi K3 the open-weight leader at 55.4 overall, with a reasoning index of 54.5 and a coding index of 45.4, ahead of GLM-5.2 in second place at 46.8.
Neither list is gospel, but together they show open-weight AI models occupying territory that was exclusively closed twelve months ago. The interesting question in late 2026 is no longer whether open weights are good enough — it is which of them you are allowed to use, and on what iron.
What “Open Weight” Really Means in 2026
Moonshot AI itself describes Kimi K3 as “open weight” rather than “open source”, a distinction Simon Willison highlighted when the weights landed. Open weights means you can download the parameters; it says nothing about training data, training code or what the licence lets you build. Some open-weight AI models arrive under clean MIT or Apache 2.0 terms; others carry bespoke conditions that only reveal their teeth at scale.
This guide therefore treats the licence as a first-class specification, not a footnote. A benchmark score tells you what a model can do; the licence tells you what you can do with the model. Both halves matter.
Who This Guide Is For
This roundup is written for CTOs, ML leads and technically minded founders comparing open-weight AI models for real deployments: coding assistants, document pipelines, on-premises agents, edge devices. If you want architecture-level detail on any single model, follow the cross-links to our dedicated articles rather than expecting a full teardown here — the point of this page is the comparison, kept honest with sourced figures.
How to Judge Open-Weight AI Models: Licence Tiers Meet Hardware Tiers
Leaderboards answer one question; deployments ask two more. Before any benchmark matters you need to know whether the licence permits your use case and whether your hardware can host the weights. We therefore sort the field along two axes and place every model in the resulting grid. It is a deliberately unglamorous way to rank open-weight AI models, and it is the one that predicts real-world adoption best.
The Three Licence Tiers of Open-Weight AI Models
Tier one is truly permissive: MIT and Apache 2.0. GLM-5.2 and DeepSeek V4-Flash ship under MIT; Gemma 4, Muse Glimmer, gpt-oss and Mistral Large 3 under Apache 2.0. Tier two is custom-but-commercial: NVIDIA’s Open Model Agreement for Nemotron 3 Nano Omni is described as commercially usable, and MiniMax M3 ships under its minimax-community licence. Tier three is revenue-gated or bespoke: Kimi K3’s custom K3 licence and the qwen3.8-max licence attached to Alibaba’s flagship weights. Among open-weight AI models, tier three demands a lawyer before a GPU.
The Four Hardware Tiers
At the bottom sits the 16GB consumer card, where gpt-oss-20b fits comfortably. Next, the 24–32GB enthusiast tier hosts Muse Glimmer (24GB at 4-bit) and Nemotron 3 Nano Omni (~25GB). The third tier is the single 80GB datacentre GPU — gpt-oss-120b in MXFP4 is the canonical example. Above that live multi-GPU nodes for the 300B–1T sparse models, and finally terabyte-class clusters for Kimi K3’s 1.56TB and Qwen3.8-Max’s roughly 4.89TB downloads. Matching open-weight AI models to the tier you actually own is step one of any evaluation.
The Decision Matrix at a Glance
The grid below crosses the two axes. It is the single most useful table in this guide: find your hardware column, drop down to the licence row your business can accept, and your shortlist of open-weight AI models writes itself.
| Licence tier | Consumer GPU (16–32GB) | Single 80GB GPU | Multi-GPU node | Terabyte cluster |
|---|---|---|---|---|
| Permissive (MIT / Apache 2.0) | gpt-oss-20b, Muse Glimmer, Gemma 4 (small sizes) | gpt-oss-120b, Gemma 4 31B | GLM-5.2, DeepSeek V4-Flash, Mistral Large 3 | — |
| Custom, commercially usable | Nemotron 3 Nano Omni | — | MiniMax M3; Inkling (licence unstated) | — |
| Revenue-gated / bespoke | — | — | — | Kimi K3 (K3 licence), Qwen3.8-Max (qwen3.8-max licence) |
The Eleven Best Open-Weight AI Models at a Glance
Eleven models earn a place in this roundup, spanning three countries, three licence tiers and hardware demands that run from a 16GB consumer card to a roughly 4.89TB download. Every specification in the master table below comes from the model card, the developer’s announcement or an independent measurement — and each of these open-weight AI models is covered in more depth later on this page.
Master Table: Eleven Open-Weight AI Models Compared
| Model | Developer | Weights released | Total / active params | Context | Licence |
|---|---|---|---|---|---|
| Kimi K3 | Moonshot AI | 27 Jul 2026 | 2.8T / ~104B | 1M tokens | Custom K3 licence |
| Qwen3.8-Max | Alibaba | 8 Aug 2026 | 2.4T / 95B | 262,144 native (to 1,010,000) | Bespoke qwen3.8-max |
| GLM-5.2 | Zhipu (Z.ai) | 13 Jun 2026 | 744B / 40B | 1M tokens | MIT |
| DeepSeek V4-Flash-0731 | DeepSeek | 31 Jul 2026 | 304B (284B base) / ~13B | 1M tokens | MIT |
| MiniMax M3 | MiniMax | 1 Jun 2026 | ~428B / ~23B | 1M tokens | minimax-community |
| Inkling | Thinking Machines Lab | 15 Jul 2026 | 975B / 41B | Up to 1M tokens | Open weights (Hugging Face) |
| Muse Glimmer | Meta | 10 Aug 2026 | 30B dense | 131,072+ tokens | Apache 2.0 |
| Gemma 4 family | 2 Apr 2026 | E2B to 31B dense | — | Apache 2.0 | |
| gpt-oss-120b / 20b | OpenAI | 5 Aug 2025 | 117B / 5.1B and 21B / 3.6B | 128K tokens | Apache 2.0 |
| Mistral Large 3 | Mistral AI | Dec 2025 | 675B / 41B | 256K tokens | Apache 2.0 |
| Nemotron 3 Nano Omni | NVIDIA | 28 Apr 2026 | 30B / 3B | — | NVIDIA Open Model Agreement |
How to Read the Table
Three patterns jump out. First, sparse mixture-of-experts is now the default architecture for large open-weight AI models — most of the field activates a small fraction of total parameters per token. Second, the 1-million-token context window has become table stakes at the top end. Third, the two largest releases are precisely the two with the most restrictive licences, which is not a coincidence: the more a release cost to train, the more carefully its owner fences the commercial upside.
Frontier-Class Open-Weight AI Models: Kimi K3 and Qwen3.8-Max
Two releases define the ceiling this year, and both come from China. They are the largest, strongest and most legally complicated open-weight AI models ever shipped, and they deserve to be read together: one optimises for headline intelligence, the other for sheer scale of release.
Kimi K3: The Highest-Ranked of All Open-Weight AI Models
Announced by Moonshot AI on 16 July 2026, Kimi K3 is a 2.8-trillion-parameter sparse MoE that activates 16 of its 896 experts per token, pairing a 1-million-token context window with a hybrid Kimi Delta Attention (KDA) and Attention Residuals architecture. The weights followed on 27 July — a 1.56TB Hugging Face download, roughly 1.4TB in MXFP4, with about 104B active parameters. Our Kimi K3 deep dive unpacks the architecture properly.
2.8T total / ~104B active parameters · 16 of 896 experts per token · 1M-token context · 1.56TB weights (≈1.4TB MXFP4) · Intelligence Index 57 (#3 overall) · API $3.00 / $15.00 per 1M tokens, cached input $0.30 · custom K3 licence
On capability, the numbers are startling for an open release: 88.3 on Terminal-Bench 2.1 and an outright lead on Program Bench, though it trails on DeepSWE and FrontierSWE. Its per-task agentic cost averages $0.94 — close to GPT-5.6 Sol’s $1.04 and roughly half the price of Opus 4.8.
The K3 Licence and the US$20 Million Clause
Kimi K3 does not ship under MIT or Apache 2.0. Its custom K3 licence requires any company running it as a Model-as-a-Service business with aggregate revenue above US$20 million over any consecutive twelve months to sign a separate agreement with Moonshot AI. For most internal deployments that clause never bites; for anyone reselling inference it is the first thing to price in. It is the clearest example yet of open-weight AI models diverging from open-source software norms.
Qwen3.8-Max: The Largest Open Weights Ever Shipped
Alibaba previewed Qwen3.8-Max at WAIC Shanghai on 19 July 2026 and took it to full general availability on 3 August. It is a 2.4-trillion-parameter sparse MoE with 95B activated parameters and a native context of 262,144 tokens, extensible to 1,010,000. The model card claims formidable scores: GPQA Diamond 92.6, SWE-bench Pro 67.7, Terminal Bench 2.1 86.6, DeepSWE 1.1 56.6 and PaperBench 93.0. Our earlier Qwen3.8-Max preview coverage traced the WAIC announcement in detail.
API pricing is reported at $2.00 input and $6.00 output per million tokens, though those figures were absent from Alibaba’s own Model Studio pricing page at the time of writing — a caveat worth remembering when budgeting.
Text-Only Weights Under a Bespoke Licence
The weights themselves went live on Hugging Face on 8 August 2026 as Qwen/Qwen3.8-2.4T-A95B plus an FP8 sibling: 224 files totalling roughly 4.89TB, public and ungated. Two asterisks apply. The open release is text-only — the API version’s image and video input is absent — and it ships under a bespoke “qwen3.8-max” licence rather than the Apache 2.0 terms earlier Qwen generations made famous. Anyone treating all Qwen releases as automatically permissive should read this licence line twice.
The Missing Qwen3.8-27B
The release Qwen-watchers most want is the one still missing. A Qwen3.8-27B checkpoint was promised alongside the flagship, sized for the single-GPU tier where most teams actually deploy open-weight AI models — yet as of 13 August 2026 it had still not appeared. Its arrival is the most keenly awaited event on the open-weights calendar right now, and any day could be the day; check the Qwen Hugging Face organisation before you commit to an alternative in that size class.
Production-Grade Open-Weight AI Models: GLM-5.2, DeepSeek V4 and MiniMax M3
Below the two giants sits the most competitive stratum of the market: models in the 300B–750B range that a well-provisioned multi-GPU node can serve, with licences a normal business can live with. For most production teams, these are the open-weight AI models that will actually end up in the stack.
GLM-5.2: The MIT-Licensed Coding Workhorse
Zhipu’s GLM-5.2 launched on 13 June 2026 with 744B total and 40B active parameters, a 1-million-token context window — roughly five times GLM-5.1’s ~200K — and a maximum output of 131,072 tokens, all under a clean MIT licence. It scores 51 on the Artificial Analysis Intelligence Index (July 2026 update) and posts SWE-bench Pro 62.1, beating GPT-5.5’s 58.6, alongside Terminal-Bench 2.1 81.0 and AIME 2026 99.2%. To see how far the line has come, compare it with what GLM-5.1 brought only months earlier.
The combination — MIT terms, top-five intelligence, strong coding — makes GLM-5.2 the default recommendation of this guide for teams wanting serious open-weight AI models without licence anxiety.
A Note on the GLM-5.5 Rumour
You may have read that GLM-5.5 is imminent. Be careful: GLM-5.5 is unreleased. An August 2026 launch was reported via a JPMorgan research note relayed by Reuters on 25 June 2026, and July community leaks describe more than 1T total parameters with a 1M context — but Zhipu has published no model card, no benchmark and no endpoint. Until any of those exist, GLM-5.2 is the real, shipping model, and this guide ranks only what has shipped.
DeepSeek V4-Flash: Open-Weight AI Models at Commodity Prices
DeepSeek previewed its V4 family on 24 April 2026 and released V4-Flash-0731 with open weights under MIT on 31 July. The design brief is efficiency: 304B total parameters with around 13B active per token — the 304B describing the GA Hugging Face checkpoint with its DSpark speculative-decoding module attached, atop a 284B base model — plus a 1M-token context and up to 384K output tokens. The API price is the story — $0.14 per million input tokens on a cache miss and $0.28 output, with cache hits discounted roughly 98%. No other entry on this page moves the cost-per-token floor so far down.
DeepSeek V4-Pro: General Availability, Weights Included
The bigger sibling, V4-Pro-0813, carries 1.6 trillion total parameters with 49B active and reached general availability on 12–13 August 2026 — and on 13 August its full weights landed on Hugging Face: roughly 1.65TB in FP8, under an MIT licence, joining the previously released V4-Flash weights.
Vendor-reported benchmarks are frontier-class: SWE-bench Verified 80.6%, GPQA Diamond 90.1%, LiveCodeBench 93.5%, MMLU-Pro 87.5% and a Codeforces rating of 3,206, with API pricing of $0.435 input and $0.87 output per million tokens. The weights arrived as this roundup went to press, so V4-Pro sits outside the eleven-model tables above — but it is now a genuinely downloadable frontier system in the terabyte-cluster tier, not merely an API with a promise attached.
MiniMax M3: Multimodal Coding Value
MiniMax M3 arrived on 1 June 2026 as an open-weight coding-focused model with vendor-claimed frontier results, including 59.0% on SWE-Bench Pro — above GPT-5.5 and Gemini 3.1 Pro — though those launch claims were unverified at the time. The model card lists ~428B total and ~23B activated parameters, a 1M context and native multimodality across text, image and video; its MiniMax Sparse Attention delivers 9x prefill and 15x decode speedups versus M2 at 1M context.
Card benchmarks include SWE-bench Verified 80.5%, MMMU Pro 78.1% and Video-MME v2 85.4%, and Artificial Analysis scores it 55. Among multimodal open-weight AI models it is arguably the best value going: $0.30/$1.20 per million tokens up to 512K context, $0.60/$2.40 beyond, under the minimax-community licence.
American Open-Weight AI Models: Inkling, Muse Glimmer and gpt-oss
For eighteen months the open-weights story was overwhelmingly Chinese. The summer of 2026 changed that, with three US labs shipping meaningfully different answers to the question of what American open-weight AI models should look like.
Inkling: A Statement Release from Thinking Machines Lab
Thinking Machines Lab released Inkling on 15 July 2026: a 975B-total, 41B-active MoE transformer with a context window of up to 1M tokens, pretrained on 45 trillion tokens of text, images, audio and video, with weights on Hugging Face and API access via Tinker. It debuts at 41 on the Artificial Analysis Intelligence Index — the leading open-weights release from a US lab, ahead of Nemotron 3 Ultra (38), Gemma 4 31B (29) and gpt-oss-120b (24).
At thinking effort 0.99 it posts AIME 2026 97.1%, GPQA Diamond 87.2%, SWE-Bench Verified 77.6%, MMMU Pro 73.5% and VoiceBench 91.4%, and a smaller sibling — Inkling-Small, at 276B with 12B active — was previewed alongside it. Our Inkling analysis looks at why its one-size-does-not-fit-all philosophy matters as much as its scores.
Muse Glimmer: Meta Returns to Apache 2.0
Meta’s last open-weight Llama release remains Llama 4 Scout and Maverick, back on 5 April 2025; Llama 4 Behemoth never shipped through H1 2026 as the company pivoted its frontier work to the Muse line. The pivot’s first open fruit is Muse Glimmer, released 10 August 2026: a 30-billion-parameter multimodal model under Apache 2.0, distilled from the larger closed Muse Spark and built for always-on local agents.
The engineering is deployment-first. At 4-bit quantisation it runs on a single 24GB-VRAM GPU with 1.0% quality degradation — the smallest build keeps its weights under 20GB, leaving room for the KV cache, perception encoder and speculative-decoding drafter — or on 32GB with 0.2%; LM Studio’s model page quotes a more conservative minimum of 26GB of total system RAM for CPU and unified-memory setups. It carries a 131,072+ token context, a ~1.8B-parameter ViT-G/14 perception encoder and a knowledge cutoff of 4 January 2026.
Benchmarks: SWE-Bench Pro 51.2, AIME 2026 94.7, MCP Atlas 75.5, Gaia2 43.3 and DeepSearch QA 74.6. For local agent builders, this is one of the most practical open-weight AI models on the list.
gpt-oss: OpenAI’s Incumbent Small Weights
OpenAI’s gpt-oss pair, released 5 August 2025, remains the US incumbent at the small end. gpt-oss-120b packs 117B total and 5.1B active parameters and fits a single 80GB GPU in MXFP4; gpt-oss-20b, at 21B total and 3.6B active, fits a 16GB consumer GPU. Both carry a 128K context and clean Apache 2.0 terms. A year on, newer open-weight AI models out-score them — Artificial Analysis has gpt-oss-120b at 24 — but for frictionless licensing on modest hardware they remain the path of least resistance.
Small, Fast and European: The Rest of the Field
Not every valuable release chases the leaderboard summit. The remaining three entrants win on breadth of sizes, raw speed and European provenance respectively — and round out the picture of where open-weight AI models are heading at the edge.
Gemma 4: Apache 2.0 Across Five Sizes
Google’s Gemma 4 launched on 2 April 2026 in five sizes — E2B, E4B, 12B, a 26B-A4B MoE and a 31B dense model — accepting text, image, audio and video input. The headline was legal, not technical: for the first time in the family’s history, every Gemma 4 size ships under Apache 2.0 rather than a custom Google licence. That single change moved the whole family into the permissive tier and made Gemma 4 the easiest multi-size fleet of open-weight AI models to standardise on.
Nemotron 3 Nano Omni: The Throughput Champion
NVIDIA’s Nemotron 3 Nano Omni (28 April 2026) is the field’s speed specialist: a 30B-total, 3B-active hybrid Mamba-Transformer MoE that natively unifies vision, audio and language, runs in about 25GB of VRAM and claims 9x higher throughput than other open omni models. Per BenchLM measurements on 5 August 2026 it was the fastest measured open model at 323 tokens per second. Unusually, NVIDIA ships weights, datasets and training recipes together under its commercially usable Open Model Agreement — a completeness few open-weight AI models match.
Mistral Large 3 and the Family in Early Access
Europe’s flagship remains Mistral Large 3 (December 2025): a 675B-total, 41B-active MoE under Apache 2.0 with a 256K context, native text-plus-image input and La Plateforme pricing of $0.50/$1.50 per million tokens. Its successor is already moving — Mistral confirmed in July 2026 that a new open-weight model family had entered early access with research, government and industry partners, with CEO Arthur Mensch declining to disclose parameter count, benchmarks or licence, and a broader release expected later in summer 2026. Until then, Large 3 holds the line for European open-weight AI models.
Benchmarks: How the Leading Open Weights Stack Up
Benchmark tables reward careful reading: vendors quote different suites, different variants of the same suite and different harness settings. Everything below is labelled with its exact source, and where a figure is vendor-claimed rather than independently measured, we say so.
Where Open-Weight AI Models Rank on the Intelligence Index
Artificial Analysis provides the cleanest single-number comparison. Its Intelligence Index scores Kimi K3 at 57, MiniMax M3 at 55, GLM-5.2 at 51 (July 2026 update), Inkling at 41, Nemotron 3 Ultra at 38, Gemma 4 31B at 29 and gpt-oss-120b at 24. The takeaway: the top open-weight AI models now cluster within a few points of each other, while the US pack trails the Chinese leaders by a visible margin.
Coding Benchmarks for Open-Weight AI Models
On SWE-bench Pro — the harder variant — the sourced figures line up as follows: Qwen3.8-Max 67.7 (model card), GLM-5.2 62.1, MiniMax M3 59.0 (vendor-claimed at launch), closed GPT-5.5 at 58.6 for reference, and Muse Glimmer 51.2. Note that SWE-bench Verified is a different, easier suite: DeepSeek V4-Pro reports 80.6% and MiniMax M3’s card 80.5% there, with Inkling at 77.6% — never compare a Pro number with a Verified number directly.
The wider benchmark picture, each figure against its exact suite:
| Model | Benchmark | Score | Status |
|---|---|---|---|
| Kimi K3 | Terminal-Bench 2.1 | 88.3 | Reported (DataCamp) |
| Qwen3.8-Max | GPQA Diamond | 92.6 | Model card |
| Qwen3.8-Max | Terminal Bench 2.1 | 86.6 | Model card |
| GLM-5.2 | AIME 2026 | 99.2% | Reported (Fello AI) |
| DeepSeek V4-Pro-0813 | SWE-bench Verified | 80.6% | Vendor-reported |
| MiniMax M3 | SWE-bench Verified | 80.5% | Model card |
| Inkling | AIME 2026 | 97.1% | Lab-reported (effort 0.99) |
| Muse Glimmer | AIME 2026 | 94.7 | Reported (MarkTechPost) |
| Nemotron 3 Nano Omni | BenchLM throughput | 323 tokens/s | Independently measured |
Speed: Tokens per Second and Long-Context Throughput
Raw intelligence is only half the operational story. Nemotron 3 Nano Omni’s independently measured 323 tokens per second (BenchLM, 5 August 2026) makes it the fastest open model tested, which matters enormously for voice and real-time agents. At the long-context end, MiniMax M3’s sparse attention posts 9x prefill and 15x decode speedups over M2 at 1M tokens — a reminder that among open-weight AI models, architecture choices now show up directly on the latency bill.
API Pricing and Running Costs
Self-hosting is not the only way to consume open weights: every major entry also has hosted API pricing, which doubles as a market signal of what serving each model really costs. The spread is enormous.
Price per Million Tokens Compared
| Model | Input $/1M | Output $/1M | Notes |
|---|---|---|---|
| Kimi K3 | $3.00 | $15.00 | Cached input $0.30 (90% off); ~$0.94 per agentic task |
| Qwen3.8-Max | $2.00 | $6.00 | Reported; absent from Model Studio pricing page |
| Mistral Large 3 | $0.50 | $1.50 | La Plateforme |
| DeepSeek V4-Pro | $0.435 | $0.87 | Cache hits discounted ~98% |
| MiniMax M3 | $0.30 | $1.20 | Up to 512K context; $0.60/$2.40 for 512K–1M |
| DeepSeek V4-Flash | $0.14 | $0.28 | Cache-miss input price; ~98% cache-hit discount |
The Output-Price Gap Visualised
The output-token column above spans a 54-fold range — from Kimi K3’s $15.00 per million down to DeepSeek V4-Flash’s $0.28 — and the chart makes the cliff between the frontier tier and the efficiency tier unmistakable.
Subscription Routes: The GLM Coding Plan
Zhipu monetises GLM-5.2 differently, through the GLM Coding Plan: Lite at $18 per month for roughly 80 prompts per five hours, Pro at $72 for around 400 and Max at $160 for about 1,600, with 20% off annual billing. For individual developers and small teams, a flat subscription against one of the strongest coding open-weight AI models is often cheaper than metered tokens — and far easier to budget.
Licensing Deep Dive: Reading the Small Print
Licence terms decide more open-weights procurement outcomes than benchmarks do. This section reads the small print across all three tiers so you do not have to discover it in a compliance review.
Truly Permissive Open-Weight AI Models: MIT and Apache 2.0
Six of our eleven entries are genuinely permissive. GLM-5.2 and DeepSeek V4-Flash use MIT; Gemma 4 (a first for that family), Muse Glimmer, gpt-oss and Mistral Large 3 use Apache 2.0. These are the open-weight AI models you can embed in commercial products, fine-tune, redistribute and resell without negotiating anything. If your product roadmap involves shipping the model itself to customers, start — and possibly end — your search in this tier.
Revenue-Gated and Bespoke Licences
The frontier tier is another world. Kimi K3’s custom K3 licence triggers a mandatory agreement with Moonshot once Model-as-a-Service revenue passes US$20 million in any twelve-month window. Qwen3.8-Max ships under its bespoke qwen3.8-max licence rather than Apache 2.0, with the open release additionally limited to text-only input. In between sit the custom-but-usable terms: NVIDIA’s commercially usable Open Model Agreement and the minimax-community licence. None of these makes the models unusable — but each makes “just download it” the wrong first step for open-weight AI models in this tier.
What the Small Print Means for Your Product
Three practical rules follow. First, classify by licence tier before benchmarking; a disqualified model’s scores are noise. Second, watch for scale triggers — a clause that is irrelevant at £1 million of revenue can be existential at £20 million. Third, remember that weights-available is not open-source: training data and code usually stay closed even for permissive open-weight AI models, which matters if your regulator asks provenance questions.
Hardware: What It Takes to Run Them Locally
The download button is free; the electricity is not. Here is what each tier of open-weight AI models actually demands from your infrastructure, using only the figures their developers publish.
Consumer GPUs: 16GB to 32GB of VRAM
Genuine laptop-and-workstation options exist at last. OpenAI’s gpt-oss-20b fits a 16GB consumer GPU. Meta’s Muse Glimmer runs at 4-bit quantisation on a 24GB-VRAM GPU with just 1.0% quality degradation, or on 32GB with 0.2% — though if you run on CPU or unified memory, budget for LM Studio’s more conservative minimum of 26GB of system RAM. NVIDIA’s Nemotron 3 Nano Omni needs about 25GB. For prototyping agents, private assistants and edge deployments, these three cover text, vision and audio between them — no cluster required.
The Single 80GB GPU Class
One professional card unlocks the next band: gpt-oss-120b fits a single 80GB GPU in MXFP4 thanks to its 5.1B active parameters. This class is the sweet spot for departmental servers — and it is precisely where the still-missing Qwen3.8-27B would land, which is why that checkpoint matters so much to teams standardising on single-GPU open-weight AI models.
Terabyte-Class Downloads
At the top, the numbers turn logistical. Kimi K3’s weights total 1.56TB on Hugging Face, roughly 1.4TB in MXFP4. Qwen3.8-Max’s release spans 224 files at about 4.89TB. Before dreaming of self-hosting either, price the storage, the inter-GPU bandwidth and the ops time — for most organisations, the hosted APIs of these frontier open-weight AI models are the rational route, with self-hosting reserved for genuine sovereignty requirements.
How to Choose the Right Model for Your Team
All the data above compresses into a short set of defaults. Treat these as starting points to be validated against your own workloads — benchmarks are a proxy, and your prompts are the only test set that counts.
Recommendations: Matching Open-Weight AI Models to Use Cases
For production coding on your own hardware, GLM-5.2 is the default: MIT licence, SWE-bench Pro 62.1, 1M context. For maximum capability regardless of licence complexity, Kimi K3 leads all open-weight AI models at Intelligence Index 57. For cost-sensitive high-volume pipelines, DeepSeek V4-Flash at $0.14/$0.28 per million tokens is untouchable. For multimodal agents on a workstation, choose between Muse Glimmer (Apache 2.0, 24GB) and Nemotron 3 Nano Omni (323 tokens per second). For a US-lab requirement, Inkling is the strongest option; for a multi-size fleet under one clean licence, Gemma 4; for European provenance, Mistral Large 3.
When to Stay on Closed APIs
Honesty demands the counter-case. The Intelligence Index still puts Claude Fable 5 and GPT-5.6 Sol above every open release, and frontier open-weight AI models carry real operational weight — terabyte downloads, licence review, serving expertise. If your usage is modest, your data unregulated and your team small, a closed API may remain the cheaper total-cost answer for another year. The right conclusion from this guide is options, not ideology.
Where Progressive Robot Can Help
Progressive Robot helps organisations move from shortlist to production: benchmarking candidate open-weight AI models against your actual workloads, building the fine-tuning and serving stack, and wiring the result into your systems. Our ML model development team runs exactly these evaluations, and the strategy practice linked earlier handles the licence-and-economics groundwork that should precede any GPU purchase.
The Road Ahead for Open Weights
August 2026 is a snapshot of a fast-moving field. Here is what the fact base says is coming — and what remains rumour.
Releases to Watch Before Year-End
Four threads to follow. The Qwen3.8-27B checkpoint, promised but still absent as of 13 August, would instantly become the single-GPU release of the year. DeepSeek’s V4-Pro weights — the one thread already resolved — landed on Hugging Face on 13 August, roughly 1.65TB in FP8 under MIT, putting a 1.6T-parameter system into open hands.
Mistral’s unnamed new family, in early access with partners since July, is expected to release more broadly later in summer 2026. And GLM-5.5 remains strictly a rumour — a JPMorgan note relayed by Reuters plus community leaks, with no model card, benchmark or endpoint. Inkling-Small, previewed at 276B with 12B active, rounds out the watchlist for open-weight AI models.
Risks and Caveats
Three cautions temper the optimism. Vendor-claimed benchmarks — MiniMax M3’s launch figures were explicitly unverified — should be re-tested before they drive decisions. Licence drift is real: Qwen moved from Apache 2.0 heritage to a bespoke licence, and others may follow as training costs climb. And openness itself can narrow — a text-only weights release beside a multimodal API, as with Qwen3.8-Max, shows vendors learning to open some capabilities while keeping the differentiating ones closed. Evaluate open-weight AI models on what is actually released, not what the launch blog implies.
Frequently Asked Questions
What Are Open-Weight AI Models?
Open-weight AI models are systems whose trained parameters — the weights — are published for anyone to download and run, typically via Hugging Face. Unlike full open-source projects, the training data and training code usually stay private, and the licence may impose commercial conditions. The term became standard precisely because vendors like Moonshot AI describe releases such as Kimi K3 as “open weight” rather than “open source”.
Are Open-Weight AI Models Free for Commercial Use?
Often, but never assume it. Six models in this guide — GLM-5.2, DeepSeek V4-Flash, Gemma 4, Muse Glimmer, gpt-oss and Mistral Large 3 — carry MIT or Apache 2.0 terms with no commercial gate. Kimi K3, by contrast, requires a separate agreement with Moonshot once Model-as-a-Service revenue exceeds US$20 million over twelve months, and Qwen3.8-Max uses its own bespoke licence. Read the licence file before the model card.
Which Open-Weight AI Models Run on a Single GPU?
On consumer hardware: gpt-oss-20b fits a 16GB card, Muse Glimmer runs on 24GB at 4-bit quantisation and Nemotron 3 Nano Omni needs about 25GB. One 80GB professional GPU hosts gpt-oss-120b in MXFP4. Everything larger — GLM-5.2, MiniMax M3, Inkling, DeepSeek V4-Flash and the two frontier giants — needs multiple GPUs or a hosted API.
How Do Open-Weight AI Models Compare with Closed Frontier Models?
Closer than ever, without full parity. Kimi K3’s Intelligence Index score of 57 sits comparable to Claude Opus 4.8 and GPT-5.5, behind only Claude Fable 5 and GPT-5.6 Sol. On specific suites, open-weight AI models already win some head-to-heads — GLM-5.2’s SWE-bench Pro 62.1 beats GPT-5.5’s 58.6. The frontier remains closed, but the moat has narrowed to single digits.
Which of the Open-Weight AI Models Is Best for Coding?
It depends on the harness. By SWE-bench Pro, Qwen3.8-Max leads at 67.7 with GLM-5.2 at 62.1. On terminal-style agentic work, Kimi K3’s 88.3 on Terminal-Bench 2.1 is the standout, and it leads Program Bench outright. For subscription-based coding on a budget, GLM-5.2 through the GLM Coding Plan at $18 per month is the pragmatic pick of the open-weight AI models covered here.
References
Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index (Artificial Analysis)
Simon Willison on the Kimi K3 licence
Qwen/Qwen3.8-2.4T-A95B — Hugging Face model card
Introducing Inkling — Thinking Machines Lab
MiniMaxAI/MiniMax-M3 — Hugging Face model card
DeepSeek V4: Release Date, Specs, and How to Access It (Yotta Labs)
Gemma 4: Byte for byte, the most capable open models (Google)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.