Open-weight AI models have finally caught the frontier. In July 2026 Moonshot AI announced Kimi K3, and within weeks a model anyone can download was sitting at #3 on the Artificial Analysis Intelligence Index with a score of 57 — comparable to Claude Opus 4.8 and GPT-5.5, behind only Claude Fable 5 and GPT-5.6 Sol. For the first time in the short history of modern artificial intelligence, downloadable weights rank beside the best closed systems in the world.

Yet “open” has never meant so many different things at once. The label now stretches from genuine MIT and Apache 2.0 releases to revenue-gated bespoke licences, and from a 1.56TB datacentre download to a 30-billion-parameter agent that lives happily on a 24GB consumer graphics card. Treating every release as interchangeable is the fastest way to pick a model your legal team will later veto or your hardware simply cannot host.

That is why this guide ranks the eleven confirmed open-weight AI models of August 2026 by what you can actually run and legally ship — a licence tier times hardware tier decision matrix — rather than by leaderboard position alone. It sits alongside the deep dives in our AI models and tools hub, which covers each release in far more detail than a roundup can.

If your organisation is deciding whether self-hosted machine learning belongs in its stack at all, our AI strategy consultants work through exactly these licence, cost and infrastructure trade-offs every week. Consider this article the map; the territory changes monthly, and every figure below is dated and sourced so you can check it has not moved.

Why Open-Weight AI Models Matter More Than Ever in 2026

open-weight AI models - best open weight ai models 2026 b three solid hexagonal slabs

The economics of frontier AI shifted decisively this summer. Open-weight AI models used to trail the closed frontier by a year or more; today the gap is measured in single-digit index points. That changes procurement conversations: when downloadable weights score within touching distance of the best proprietary systems, paying a premium for a closed API needs a justification beyond raw capability.

Control is the second driver. Open-weight AI models can be fine-tuned on private data, pinned to a version forever, audited for behaviour and run inside your own security perimeter. None of that is possible when a vendor can retire or retrain your model overnight. For regulated industries and sovereignty-minded governments, that difference is often decisive on its own.

The Month Open-Weight AI Models Caught the Frontier

Two independent measurements anchor the claim. Artificial Analysis places Kimi K3 at 57 on its Intelligence Index — third overall, comparable to Claude Opus 4.8 and GPT-5.5. Separately, the LLM Stats Open LLM Leaderboard of 7 August 2026 makes Kimi K3 the open-weight leader at 55.4 overall, with a reasoning index of 54.5 and a coding index of 45.4, ahead of GLM-5.2 in second place at 46.8.

Neither list is gospel, but together they show open-weight AI models occupying territory that was exclusively closed twelve months ago. The interesting question in late 2026 is no longer whether open weights are good enough — it is which of them you are allowed to use, and on what iron.

What “Open Weight” Really Means in 2026

Moonshot AI itself describes Kimi K3 as “open weight” rather than “open source”, a distinction Simon Willison highlighted when the weights landed. Open weights means you can download the parameters; it says nothing about training data, training code or what the licence lets you build. Some open-weight AI models arrive under clean MIT or Apache 2.0 terms; others carry bespoke conditions that only reveal their teeth at scale.

This guide therefore treats the licence as a first-class specification, not a footnote. A benchmark score tells you what a model can do; the licence tells you what you can do with the model. Both halves matter.

Who This Guide Is For

This roundup is written for CTOs, ML leads and technically minded founders comparing open-weight AI models for real deployments: coding assistants, document pipelines, on-premises agents, edge devices. If you want architecture-level detail on any single model, follow the cross-links to our dedicated articles rather than expecting a full teardown here — the point of this page is the comparison, kept honest with sourced figures.

How to Judge Open-Weight AI Models: Licence Tiers Meet Hardware Tiers

best open weight ai models 2026 c wide mouth funnel

Leaderboards answer one question; deployments ask two more. Before any benchmark matters you need to know whether the licence permits your use case and whether your hardware can host the weights. We therefore sort the field along two axes and place every model in the resulting grid. It is a deliberately unglamorous way to rank open-weight AI models, and it is the one that predicts real-world adoption best.

The Three Licence Tiers of Open-Weight AI Models

Tier one is truly permissive: MIT and Apache 2.0. GLM-5.2 and DeepSeek V4-Flash ship under MIT; Gemma 4, Muse Glimmer, gpt-oss and Mistral Large 3 under Apache 2.0. Tier two is custom-but-commercial: NVIDIA’s Open Model Agreement for Nemotron 3 Nano Omni is described as commercially usable, and MiniMax M3 ships under its minimax-community licence. Tier three is revenue-gated or bespoke: Kimi K3’s custom K3 licence and the qwen3.8-max licence attached to Alibaba’s flagship weights. Among open-weight AI models, tier three demands a lawyer before a GPU.

The Four Hardware Tiers

At the bottom sits the 16GB consumer card, where gpt-oss-20b fits comfortably. Next, the 24–32GB enthusiast tier hosts Muse Glimmer (24GB at 4-bit) and Nemotron 3 Nano Omni (~25GB). The third tier is the single 80GB datacentre GPU — gpt-oss-120b in MXFP4 is the canonical example. Above that live multi-GPU nodes for the 300B–1T sparse models, and finally terabyte-class clusters for Kimi K3’s 1.56TB and Qwen3.8-Max’s roughly 4.89TB downloads. Matching open-weight AI models to the tier you actually own is step one of any evaluation.

The Decision Matrix at a Glance

The grid below crosses the two axes. It is the single most useful table in this guide: find your hardware column, drop down to the licence row your business can accept, and your shortlist of open-weight AI models writes itself.

Licence tierConsumer GPU (16–32GB)Single 80GB GPUMulti-GPU nodeTerabyte cluster
Permissive (MIT / Apache 2.0)gpt-oss-20b, Muse Glimmer, Gemma 4 (small sizes)gpt-oss-120b, Gemma 4 31BGLM-5.2, DeepSeek V4-Flash, Mistral Large 3
Custom, commercially usableNemotron 3 Nano OmniMiniMax M3; Inkling (licence unstated)
Revenue-gated / bespokeKimi K3 (K3 licence), Qwen3.8-Max (qwen3.8-max licence)

The Eleven Best Open-Weight AI Models at a Glance

best open weight ai models 2026 d tall stack blank paper sheets

Eleven models earn a place in this roundup, spanning three countries, three licence tiers and hardware demands that run from a 16GB consumer card to a roughly 4.89TB download. Every specification in the master table below comes from the model card, the developer’s announcement or an independent measurement — and each of these open-weight AI models is covered in more depth later on this page.

Master Table: Eleven Open-Weight AI Models Compared

ModelDeveloperWeights releasedTotal / active paramsContextLicence
Kimi K3Moonshot AI27 Jul 20262.8T / ~104B1M tokensCustom K3 licence
Qwen3.8-MaxAlibaba8 Aug 20262.4T / 95B262,144 native (to 1,010,000)Bespoke qwen3.8-max
GLM-5.2Zhipu (Z.ai)13 Jun 2026744B / 40B1M tokensMIT
DeepSeek V4-Flash-0731DeepSeek31 Jul 2026304B (284B base) / ~13B1M tokensMIT
MiniMax M3MiniMax1 Jun 2026~428B / ~23B1M tokensminimax-community
InklingThinking Machines Lab15 Jul 2026975B / 41BUp to 1M tokensOpen weights (Hugging Face)
Muse GlimmerMeta10 Aug 202630B dense131,072+ tokensApache 2.0
Gemma 4 familyGoogle2 Apr 2026E2B to 31B denseApache 2.0
gpt-oss-120b / 20bOpenAI5 Aug 2025117B / 5.1B and 21B / 3.6B128K tokensApache 2.0
Mistral Large 3Mistral AIDec 2025675B / 41B256K tokensApache 2.0
Nemotron 3 Nano OmniNVIDIA28 Apr 202630B / 3BNVIDIA Open Model Agreement

How to Read the Table

Three patterns jump out. First, sparse mixture-of-experts is now the default architecture for large open-weight AI models — most of the field activates a small fraction of total parameters per token. Second, the 1-million-token context window has become table stakes at the top end. Third, the two largest releases are precisely the two with the most restrictive licences, which is not a coincidence: the more a release cost to train, the more carefully its owner fences the commercial upside.

Frontier-Class Open-Weight AI Models: Kimi K3 and Qwen3.8-Max

best open weight ai models 2026 e three identical upright cylinders

Two releases define the ceiling this year, and both come from China. They are the largest, strongest and most legally complicated open-weight AI models ever shipped, and they deserve to be read together: one optimises for headline intelligence, the other for sheer scale of release.

Kimi K3: The Highest-Ranked of All Open-Weight AI Models

Announced by Moonshot AI on 16 July 2026, Kimi K3 is a 2.8-trillion-parameter sparse MoE that activates 16 of its 896 experts per token, pairing a 1-million-token context window with a hybrid Kimi Delta Attention (KDA) and Attention Residuals architecture. The weights followed on 27 July — a 1.56TB Hugging Face download, roughly 1.4TB in MXFP4, with about 104B active parameters. Our Kimi K3 deep dive unpacks the architecture properly.

Kimi K3 at a glance

2.8T total / ~104B active parameters · 16 of 896 experts per token · 1M-token context · 1.56TB weights (≈1.4TB MXFP4) · Intelligence Index 57 (#3 overall) · API $3.00 / $15.00 per 1M tokens, cached input $0.30 · custom K3 licence

On capability, the numbers are startling for an open release: 88.3 on Terminal-Bench 2.1 and an outright lead on Program Bench, though it trails on DeepSWE and FrontierSWE. Its per-task agentic cost averages $0.94 — close to GPT-5.6 Sol’s $1.04 and roughly half the price of Opus 4.8.

The K3 Licence and the US$20 Million Clause

Kimi K3 does not ship under MIT or Apache 2.0. Its custom K3 licence requires any company running it as a Model-as-a-Service business with aggregate revenue above US$20 million over any consecutive twelve months to sign a separate agreement with Moonshot AI. For most internal deployments that clause never bites; for anyone reselling inference it is the first thing to price in. It is the clearest example yet of open-weight AI models diverging from open-source software norms.

Qwen3.8-Max: The Largest Open Weights Ever Shipped

Alibaba previewed Qwen3.8-Max at WAIC Shanghai on 19 July 2026 and took it to full general availability on 3 August. It is a 2.4-trillion-parameter sparse MoE with 95B activated parameters and a native context of 262,144 tokens, extensible to 1,010,000. The model card claims formidable scores: GPQA Diamond 92.6, SWE-bench Pro 67.7, Terminal Bench 2.1 86.6, DeepSWE 1.1 56.6 and PaperBench 93.0. Our earlier Qwen3.8-Max preview coverage traced the WAIC announcement in detail.

API pricing is reported at $2.00 input and $6.00 output per million tokens, though those figures were absent from Alibaba’s own Model Studio pricing page at the time of writing — a caveat worth remembering when budgeting.

Text-Only Weights Under a Bespoke Licence

The weights themselves went live on Hugging Face on 8 August 2026 as Qwen/Qwen3.8-2.4T-A95B plus an FP8 sibling: 224 files totalling roughly 4.89TB, public and ungated. Two asterisks apply. The open release is text-only — the API version’s image and video input is absent — and it ships under a bespoke “qwen3.8-max” licence rather than the Apache 2.0 terms earlier Qwen generations made famous. Anyone treating all Qwen releases as automatically permissive should read this licence line twice.

The Missing Qwen3.8-27B

The release Qwen-watchers most want is the one still missing. A Qwen3.8-27B checkpoint was promised alongside the flagship, sized for the single-GPU tier where most teams actually deploy open-weight AI models — yet as of 13 August 2026 it had still not appeared. Its arrival is the most keenly awaited event on the open-weights calendar right now, and any day could be the day; check the Qwen Hugging Face organisation before you commit to an alternative in that size class.

Production-Grade Open-Weight AI Models: GLM-5.2, DeepSeek V4 and MiniMax M3

best open weight ai models 2026 f single solid cube

Below the two giants sits the most competitive stratum of the market: models in the 300B–750B range that a well-provisioned multi-GPU node can serve, with licences a normal business can live with. For most production teams, these are the open-weight AI models that will actually end up in the stack.

GLM-5.2: The MIT-Licensed Coding Workhorse

Zhipu’s GLM-5.2 launched on 13 June 2026 with 744B total and 40B active parameters, a 1-million-token context window — roughly five times GLM-5.1’s ~200K — and a maximum output of 131,072 tokens, all under a clean MIT licence. It scores 51 on the Artificial Analysis Intelligence Index (July 2026 update) and posts SWE-bench Pro 62.1, beating GPT-5.5’s 58.6, alongside Terminal-Bench 2.1 81.0 and AIME 2026 99.2%. To see how far the line has come, compare it with what GLM-5.1 brought only months earlier.

The combination — MIT terms, top-five intelligence, strong coding — makes GLM-5.2 the default recommendation of this guide for teams wanting serious open-weight AI models without licence anxiety.

A Note on the GLM-5.5 Rumour

You may have read that GLM-5.5 is imminent. Be careful: GLM-5.5 is unreleased. An August 2026 launch was reported via a JPMorgan research note relayed by Reuters on 25 June 2026, and July community leaks describe more than 1T total parameters with a 1M context — but Zhipu has published no model card, no benchmark and no endpoint. Until any of those exist, GLM-5.2 is the real, shipping model, and this guide ranks only what has shipped.

DeepSeek V4-Flash: Open-Weight AI Models at Commodity Prices

DeepSeek previewed its V4 family on 24 April 2026 and released V4-Flash-0731 with open weights under MIT on 31 July. The design brief is efficiency: 304B total parameters with around 13B active per token — the 304B describing the GA Hugging Face checkpoint with its DSpark speculative-decoding module attached, atop a 284B base model — plus a 1M-token context and up to 384K output tokens. The API price is the story — $0.14 per million input tokens on a cache miss and $0.28 output, with cache hits discounted roughly 98%. No other entry on this page moves the cost-per-token floor so far down.

DeepSeek V4-Pro: General Availability, Weights Included

The bigger sibling, V4-Pro-0813, carries 1.6 trillion total parameters with 49B active and reached general availability on 12–13 August 2026 — and on 13 August its full weights landed on Hugging Face: roughly 1.65TB in FP8, under an MIT licence, joining the previously released V4-Flash weights.

Vendor-reported benchmarks are frontier-class: SWE-bench Verified 80.6%, GPQA Diamond 90.1%, LiveCodeBench 93.5%, MMLU-Pro 87.5% and a Codeforces rating of 3,206, with API pricing of $0.435 input and $0.87 output per million tokens. The weights arrived as this roundup went to press, so V4-Pro sits outside the eleven-model tables above — but it is now a genuinely downloadable frontier system in the terabyte-cluster tier, not merely an API with a promise attached.

MiniMax M3: Multimodal Coding Value

MiniMax M3 arrived on 1 June 2026 as an open-weight coding-focused model with vendor-claimed frontier results, including 59.0% on SWE-Bench Pro — above GPT-5.5 and Gemini 3.1 Pro — though those launch claims were unverified at the time. The model card lists ~428B total and ~23B activated parameters, a 1M context and native multimodality across text, image and video; its MiniMax Sparse Attention delivers 9x prefill and 15x decode speedups versus M2 at 1M context.

Card benchmarks include SWE-bench Verified 80.5%, MMMU Pro 78.1% and Video-MME v2 85.4%, and Artificial Analysis scores it 55. Among multimodal open-weight AI models it is arguably the best value going: $0.30/$1.20 per million tokens up to 512K context, $0.60/$2.40 beyond, under the minimax-community licence.

American Open-Weight AI Models: Inkling, Muse Glimmer and gpt-oss

For eighteen months the open-weights story was overwhelmingly Chinese. The summer of 2026 changed that, with three US labs shipping meaningfully different answers to the question of what American open-weight AI models should look like.

Inkling: A Statement Release from Thinking Machines Lab

Thinking Machines Lab released Inkling on 15 July 2026: a 975B-total, 41B-active MoE transformer with a context window of up to 1M tokens, pretrained on 45 trillion tokens of text, images, audio and video, with weights on Hugging Face and API access via Tinker. It debuts at 41 on the Artificial Analysis Intelligence Index — the leading open-weights release from a US lab, ahead of Nemotron 3 Ultra (38), Gemma 4 31B (29) and gpt-oss-120b (24).

At thinking effort 0.99 it posts AIME 2026 97.1%, GPQA Diamond 87.2%, SWE-Bench Verified 77.6%, MMMU Pro 73.5% and VoiceBench 91.4%, and a smaller sibling — Inkling-Small, at 276B with 12B active — was previewed alongside it. Our Inkling analysis looks at why its one-size-does-not-fit-all philosophy matters as much as its scores.

Muse Glimmer: Meta Returns to Apache 2.0

Meta’s last open-weight Llama release remains Llama 4 Scout and Maverick, back on 5 April 2025; Llama 4 Behemoth never shipped through H1 2026 as the company pivoted its frontier work to the Muse line. The pivot’s first open fruit is Muse Glimmer, released 10 August 2026: a 30-billion-parameter multimodal model under Apache 2.0, distilled from the larger closed Muse Spark and built for always-on local agents.

The engineering is deployment-first. At 4-bit quantisation it runs on a single 24GB-VRAM GPU with 1.0% quality degradation — the smallest build keeps its weights under 20GB, leaving room for the KV cache, perception encoder and speculative-decoding drafter — or on 32GB with 0.2%; LM Studio’s model page quotes a more conservative minimum of 26GB of total system RAM for CPU and unified-memory setups. It carries a 131,072+ token context, a ~1.8B-parameter ViT-G/14 perception encoder and a knowledge cutoff of 4 January 2026.

Benchmarks: SWE-Bench Pro 51.2, AIME 2026 94.7, MCP Atlas 75.5, Gaia2 43.3 and DeepSearch QA 74.6. For local agent builders, this is one of the most practical open-weight AI models on the list.

gpt-oss: OpenAI’s Incumbent Small Weights

OpenAI’s gpt-oss pair, released 5 August 2025, remains the US incumbent at the small end. gpt-oss-120b packs 117B total and 5.1B active parameters and fits a single 80GB GPU in MXFP4; gpt-oss-20b, at 21B total and 3.6B active, fits a 16GB consumer GPU. Both carry a 128K context and clean Apache 2.0 terms. A year on, newer open-weight AI models out-score them — Artificial Analysis has gpt-oss-120b at 24 — but for frictionless licensing on modest hardware they remain the path of least resistance.

Small, Fast and European: The Rest of the Field

Not every valuable release chases the leaderboard summit. The remaining three entrants win on breadth of sizes, raw speed and European provenance respectively — and round out the picture of where open-weight AI models are heading at the edge.

Gemma 4: Apache 2.0 Across Five Sizes

Google’s Gemma 4 launched on 2 April 2026 in five sizes — E2B, E4B, 12B, a 26B-A4B MoE and a 31B dense model — accepting text, image, audio and video input. The headline was legal, not technical: for the first time in the family’s history, every Gemma 4 size ships under Apache 2.0 rather than a custom Google licence. That single change moved the whole family into the permissive tier and made Gemma 4 the easiest multi-size fleet of open-weight AI models to standardise on.

Nemotron 3 Nano Omni: The Throughput Champion

NVIDIA’s Nemotron 3 Nano Omni (28 April 2026) is the field’s speed specialist: a 30B-total, 3B-active hybrid Mamba-Transformer MoE that natively unifies vision, audio and language, runs in about 25GB of VRAM and claims 9x higher throughput than other open omni models. Per BenchLM measurements on 5 August 2026 it was the fastest measured open model at 323 tokens per second. Unusually, NVIDIA ships weights, datasets and training recipes together under its commercially usable Open Model Agreement — a completeness few open-weight AI models match.

Mistral Large 3 and the Family in Early Access

Europe’s flagship remains Mistral Large 3 (December 2025): a 675B-total, 41B-active MoE under Apache 2.0 with a 256K context, native text-plus-image input and La Plateforme pricing of $0.50/$1.50 per million tokens. Its successor is already moving — Mistral confirmed in July 2026 that a new open-weight model family had entered early access with research, government and industry partners, with CEO Arthur Mensch declining to disclose parameter count, benchmarks or licence, and a broader release expected later in summer 2026. Until then, Large 3 holds the line for European open-weight AI models.

Benchmarks: How the Leading Open Weights Stack Up

Benchmark tables reward careful reading: vendors quote different suites, different variants of the same suite and different harness settings. Everything below is labelled with its exact source, and where a figure is vendor-claimed rather than independently measured, we say so.

Where Open-Weight AI Models Rank on the Intelligence Index

Artificial Analysis provides the cleanest single-number comparison. Its Intelligence Index scores Kimi K3 at 57, MiniMax M3 at 55, GLM-5.2 at 51 (July 2026 update), Inkling at 41, Nemotron 3 Ultra at 38, Gemma 4 31B at 29 and gpt-oss-120b at 24. The takeaway: the top open-weight AI models now cluster within a few points of each other, while the US pack trails the Chinese leaders by a visible margin.

Artificial Analysis Intelligence Index — open-weight models (Aug 2026)
Kimi K3 57
MiniMax M3 55
GLM-5.2 51
Inkling 41
Nemotron 3 Ultra 38
Gemma 4 31B 29
gpt-oss-120b 24

Coding Benchmarks for Open-Weight AI Models

On SWE-bench Pro — the harder variant — the sourced figures line up as follows: Qwen3.8-Max 67.7 (model card), GLM-5.2 62.1, MiniMax M3 59.0 (vendor-claimed at launch), closed GPT-5.5 at 58.6 for reference, and Muse Glimmer 51.2. Note that SWE-bench Verified is a different, easier suite: DeepSeek V4-Pro reports 80.6% and MiniMax M3’s card 80.5% there, with Inkling at 77.6% — never compare a Pro number with a Verified number directly.

SWE-bench Pro — sourced scores (%)
Qwen3.8-Max 67.7
GLM-5.2 62.1
MiniMax M3 (vendor-claimed) 59.0
GPT-5.5 (closed, reference) 58.6
Muse Glimmer 51.2

The wider benchmark picture, each figure against its exact suite:

ModelBenchmarkScoreStatus
Kimi K3Terminal-Bench 2.188.3Reported (DataCamp)
Qwen3.8-MaxGPQA Diamond92.6Model card
Qwen3.8-MaxTerminal Bench 2.186.6Model card
GLM-5.2AIME 202699.2%Reported (Fello AI)
DeepSeek V4-Pro-0813SWE-bench Verified80.6%Vendor-reported
MiniMax M3SWE-bench Verified80.5%Model card
InklingAIME 202697.1%Lab-reported (effort 0.99)
Muse GlimmerAIME 202694.7Reported (MarkTechPost)
Nemotron 3 Nano OmniBenchLM throughput323 tokens/sIndependently measured

Speed: Tokens per Second and Long-Context Throughput

Raw intelligence is only half the operational story. Nemotron 3 Nano Omni’s independently measured 323 tokens per second (BenchLM, 5 August 2026) makes it the fastest open model tested, which matters enormously for voice and real-time agents. At the long-context end, MiniMax M3’s sparse attention posts 9x prefill and 15x decode speedups over M2 at 1M tokens — a reminder that among open-weight AI models, architecture choices now show up directly on the latency bill.

API Pricing and Running Costs

Self-hosting is not the only way to consume open weights: every major entry also has hosted API pricing, which doubles as a market signal of what serving each model really costs. The spread is enormous.

Price per Million Tokens Compared

ModelInput $/1MOutput $/1MNotes
Kimi K3$3.00$15.00Cached input $0.30 (90% off); ~$0.94 per agentic task
Qwen3.8-Max$2.00$6.00Reported; absent from Model Studio pricing page
Mistral Large 3$0.50$1.50La Plateforme
DeepSeek V4-Pro$0.435$0.87Cache hits discounted ~98%
MiniMax M3$0.30$1.20Up to 512K context; $0.60/$2.40 for 512K–1M
DeepSeek V4-Flash$0.14$0.28Cache-miss input price; ~98% cache-hit discount

The Output-Price Gap Visualised

The output-token column above spans a 54-fold range — from Kimi K3’s $15.00 per million down to DeepSeek V4-Flash’s $0.28 — and the chart makes the cliff between the frontier tier and the efficiency tier unmistakable.

Output price, $ per 1M tokens (hosted APIs)
Kimi K3 $15.00
Qwen3.8-Max $6.00
Mistral Large 3 $1.50
MiniMax M3 $1.20
DeepSeek V4-Pro $0.87
DeepSeek V4-Flash $0.28

Subscription Routes: The GLM Coding Plan

Zhipu monetises GLM-5.2 differently, through the GLM Coding Plan: Lite at $18 per month for roughly 80 prompts per five hours, Pro at $72 for around 400 and Max at $160 for about 1,600, with 20% off annual billing. For individual developers and small teams, a flat subscription against one of the strongest coding open-weight AI models is often cheaper than metered tokens — and far easier to budget.

Licensing Deep Dive: Reading the Small Print

Licence terms decide more open-weights procurement outcomes than benchmarks do. This section reads the small print across all three tiers so you do not have to discover it in a compliance review.

Truly Permissive Open-Weight AI Models: MIT and Apache 2.0

Six of our eleven entries are genuinely permissive. GLM-5.2 and DeepSeek V4-Flash use MIT; Gemma 4 (a first for that family), Muse Glimmer, gpt-oss and Mistral Large 3 use Apache 2.0. These are the open-weight AI models you can embed in commercial products, fine-tune, redistribute and resell without negotiating anything. If your product roadmap involves shipping the model itself to customers, start — and possibly end — your search in this tier.

Revenue-Gated and Bespoke Licences

The frontier tier is another world. Kimi K3’s custom K3 licence triggers a mandatory agreement with Moonshot once Model-as-a-Service revenue passes US$20 million in any twelve-month window. Qwen3.8-Max ships under its bespoke qwen3.8-max licence rather than Apache 2.0, with the open release additionally limited to text-only input. In between sit the custom-but-usable terms: NVIDIA’s commercially usable Open Model Agreement and the minimax-community licence. None of these makes the models unusable — but each makes “just download it” the wrong first step for open-weight AI models in this tier.

What the Small Print Means for Your Product

Three practical rules follow. First, classify by licence tier before benchmarking; a disqualified model’s scores are noise. Second, watch for scale triggers — a clause that is irrelevant at £1 million of revenue can be existential at £20 million. Third, remember that weights-available is not open-source: training data and code usually stay closed even for permissive open-weight AI models, which matters if your regulator asks provenance questions.

Hardware: What It Takes to Run Them Locally

The download button is free; the electricity is not. Here is what each tier of open-weight AI models actually demands from your infrastructure, using only the figures their developers publish.

Consumer GPUs: 16GB to 32GB of VRAM

Genuine laptop-and-workstation options exist at last. OpenAI’s gpt-oss-20b fits a 16GB consumer GPU. Meta’s Muse Glimmer runs at 4-bit quantisation on a 24GB-VRAM GPU with just 1.0% quality degradation, or on 32GB with 0.2% — though if you run on CPU or unified memory, budget for LM Studio’s more conservative minimum of 26GB of system RAM. NVIDIA’s Nemotron 3 Nano Omni needs about 25GB. For prototyping agents, private assistants and edge deployments, these three cover text, vision and audio between them — no cluster required.

The Single 80GB GPU Class

One professional card unlocks the next band: gpt-oss-120b fits a single 80GB GPU in MXFP4 thanks to its 5.1B active parameters. This class is the sweet spot for departmental servers — and it is precisely where the still-missing Qwen3.8-27B would land, which is why that checkpoint matters so much to teams standardising on single-GPU open-weight AI models.

Terabyte-Class Downloads

At the top, the numbers turn logistical. Kimi K3’s weights total 1.56TB on Hugging Face, roughly 1.4TB in MXFP4. Qwen3.8-Max’s release spans 224 files at about 4.89TB. Before dreaming of self-hosting either, price the storage, the inter-GPU bandwidth and the ops time — for most organisations, the hosted APIs of these frontier open-weight AI models are the rational route, with self-hosting reserved for genuine sovereignty requirements.

How to Choose the Right Model for Your Team

All the data above compresses into a short set of defaults. Treat these as starting points to be validated against your own workloads — benchmarks are a proxy, and your prompts are the only test set that counts.

Recommendations: Matching Open-Weight AI Models to Use Cases

For production coding on your own hardware, GLM-5.2 is the default: MIT licence, SWE-bench Pro 62.1, 1M context. For maximum capability regardless of licence complexity, Kimi K3 leads all open-weight AI models at Intelligence Index 57. For cost-sensitive high-volume pipelines, DeepSeek V4-Flash at $0.14/$0.28 per million tokens is untouchable. For multimodal agents on a workstation, choose between Muse Glimmer (Apache 2.0, 24GB) and Nemotron 3 Nano Omni (323 tokens per second). For a US-lab requirement, Inkling is the strongest option; for a multi-size fleet under one clean licence, Gemma 4; for European provenance, Mistral Large 3.

When to Stay on Closed APIs

Honesty demands the counter-case. The Intelligence Index still puts Claude Fable 5 and GPT-5.6 Sol above every open release, and frontier open-weight AI models carry real operational weight — terabyte downloads, licence review, serving expertise. If your usage is modest, your data unregulated and your team small, a closed API may remain the cheaper total-cost answer for another year. The right conclusion from this guide is options, not ideology.

Where Progressive Robot Can Help

Progressive Robot helps organisations move from shortlist to production: benchmarking candidate open-weight AI models against your actual workloads, building the fine-tuning and serving stack, and wiring the result into your systems. Our ML model development team runs exactly these evaluations, and the strategy practice linked earlier handles the licence-and-economics groundwork that should precede any GPU purchase.

The Road Ahead for Open Weights

August 2026 is a snapshot of a fast-moving field. Here is what the fact base says is coming — and what remains rumour.

Releases to Watch Before Year-End

Four threads to follow. The Qwen3.8-27B checkpoint, promised but still absent as of 13 August, would instantly become the single-GPU release of the year. DeepSeek’s V4-Pro weights — the one thread already resolved — landed on Hugging Face on 13 August, roughly 1.65TB in FP8 under MIT, putting a 1.6T-parameter system into open hands.

Mistral’s unnamed new family, in early access with partners since July, is expected to release more broadly later in summer 2026. And GLM-5.5 remains strictly a rumour — a JPMorgan note relayed by Reuters plus community leaks, with no model card, benchmark or endpoint. Inkling-Small, previewed at 276B with 12B active, rounds out the watchlist for open-weight AI models.

Risks and Caveats

Three cautions temper the optimism. Vendor-claimed benchmarks — MiniMax M3’s launch figures were explicitly unverified — should be re-tested before they drive decisions. Licence drift is real: Qwen moved from Apache 2.0 heritage to a bespoke licence, and others may follow as training costs climb. And openness itself can narrow — a text-only weights release beside a multimodal API, as with Qwen3.8-Max, shows vendors learning to open some capabilities while keeping the differentiating ones closed. Evaluate open-weight AI models on what is actually released, not what the launch blog implies.

Frequently Asked Questions

What Are Open-Weight AI Models?

Open-weight AI models are systems whose trained parameters — the weights — are published for anyone to download and run, typically via Hugging Face. Unlike full open-source projects, the training data and training code usually stay private, and the licence may impose commercial conditions. The term became standard precisely because vendors like Moonshot AI describe releases such as Kimi K3 as “open weight” rather than “open source”.

Are Open-Weight AI Models Free for Commercial Use?

Often, but never assume it. Six models in this guide — GLM-5.2, DeepSeek V4-Flash, Gemma 4, Muse Glimmer, gpt-oss and Mistral Large 3 — carry MIT or Apache 2.0 terms with no commercial gate. Kimi K3, by contrast, requires a separate agreement with Moonshot once Model-as-a-Service revenue exceeds US$20 million over twelve months, and Qwen3.8-Max uses its own bespoke licence. Read the licence file before the model card.

Which Open-Weight AI Models Run on a Single GPU?

On consumer hardware: gpt-oss-20b fits a 16GB card, Muse Glimmer runs on 24GB at 4-bit quantisation and Nemotron 3 Nano Omni needs about 25GB. One 80GB professional GPU hosts gpt-oss-120b in MXFP4. Everything larger — GLM-5.2, MiniMax M3, Inkling, DeepSeek V4-Flash and the two frontier giants — needs multiple GPUs or a hosted API.

How Do Open-Weight AI Models Compare with Closed Frontier Models?

Closer than ever, without full parity. Kimi K3’s Intelligence Index score of 57 sits comparable to Claude Opus 4.8 and GPT-5.5, behind only Claude Fable 5 and GPT-5.6 Sol. On specific suites, open-weight AI models already win some head-to-heads — GLM-5.2’s SWE-bench Pro 62.1 beats GPT-5.5’s 58.6. The frontier remains closed, but the moat has narrowed to single digits.

Which of the Open-Weight AI Models Is Best for Coding?

It depends on the harness. By SWE-bench Pro, Qwen3.8-Max leads at 67.7 with GLM-5.2 at 62.1. On terminal-style agentic work, Kimi K3’s 88.3 on Terminal-Bench 2.1 is the standout, and it leads Program Bench outright. For subscription-based coding on a budget, GLM-5.2 through the GLM Coding Plan at $18 per month is the pragmatic pick of the open-weight AI models covered here.

References