Reflection Beam is the first model from Reflection AI, the Brooklyn-based lab that has raised billions of dollars on a promise to build America’s answer to DeepSeek and Qwen. Announced on 5 October 2026, it is a 501-billion-parameter mixture-of-experts model that activates 23 billion parameters per token, with a 1 million token context window. Reflection says it matches leading Chinese open models on advanced reasoning while using “3-4x less inference compute”, and that the weights will be released under the Apache 2.0 licence later this month.

The launch, reported by TechCrunch after an Axios scoop at the weekend, matters for anyone choosing between closed APIs and models they can run themselves. Reflection published a detailed blog post with a 21-row benchmark table, which makes it possible to check the headline claims against the company’s own numbers.

We did that, row by row. Below we cover what Reflection Beam is, where it wins and loses against each rival in Reflection’s own table, how the compute claim is calculated, how the model was trained, what it would take to run it yourself, and what is still missing before anyone can verify it.

Reflection Beam at a Glance

reflection beam open weight model chinese rivals b rowing machine with a fan flywheel

Reflection describes Beam as “a workhorse model” for coding, reasoning and agentic work, aimed at enterprises, the public sector and developers. It is text-only.

The specification

As with any large language model launch, the first numbers to check are size, data and context. The table gathers the figures Reflection has published so far.

ItemReflection Beam
ArchitectureSparse mixture-of-experts, 52 layers
Total parameters501 billion
Active parameters per token23 billion
Pretraining corpus23.8 trillion tokens (web, public and licensed)
Context window1 million tokens
ModalityText only
LicenceApache 2.0 (weights due later in October)
Availability todayWaitlist for a select group of users

How it compares on size

For context, Z.ai’s GLM-5.2, the Chinese model Reflection measures itself against most directly, has roughly 744 billion total parameters with 40 billion active. Reflection Beam is about two-thirds of that size in total parameters (501 divided by 744 is 0.67) and uses a little over half the active parameters per token (23 divided by 40 is 0.58).

Who is behind it

Reflection was founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou. According to PitchBook figures cited by TechCrunch, it has raised about $4.7 billion from backers including Nvidia, Sequoia Capital and Lightspeed, and its last round valued it at $25 billion before the new money.

Reflection Beam's Benchmarks, Row by Row

reflection beam open weight model chinese rivals c pinewood derby ramp with two cars neck and neck

Reflection’s launch post compares Beam with seven other open models on 21 benchmarks covering coding, reasoning, tool use and general skills. Not every model has a score on every test (missing scores are marked “NR”), so we counted only the rows where both Reflection Beam and the rival have a number.

The head-to-head count

Higher is better on every benchmark in the table. Here is how often Reflection Beam comes out ahead.

Rival modelShared benchmarksBeam aheadBeam behindTied
Inkling (Thinking Machines)151320
Nemotron 3 Ultra (Nvidia)151311
GLM 5.2 (Z.ai)13850
DeepSeek V4.1 Flash9270
Kimi K3 (Moonshot)141130
GLM 5.3 (Z.ai)120120
Qwen 3.8 Max (Alibaba)130130

The chart turns those counts into a share of shared benchmarks won by Reflection Beam (wins divided by shared rows).

Share of shared benchmarks where Reflection Beam scores higher (Reflection’s own table)
vs Inkling, 13 of 15 86.7%
vs Nemotron 3 Ultra, 13 of 15 86.7%
vs GLM 5.2, 8 of 13 61.5%
vs DeepSeek V4.1 Flash, 2 of 9 22.2%
vs Kimi K3, 1 of 14 7.1%
vs GLM 5.3, 0 of 12 0%
vs Qwen 3.8 Max, 0 of 13 0%

Against the Western open models

Here the claim holds up well. Reflection Beam beats Thinking Machines’ Inkling on 13 of 15 shared tests and Nvidia’s Nemotron 3 Ultra on 13 of 15. Inkling, released in July, is multimodal and Beam is not, so the comparison flatters Beam slightly, but on coding the gap is clear: SWE Bench Pro v2-Hard 77.2 against 56.9, and Terminal Bench v2.1 80.1 against 63.8.

Against GLM 5.2

This is the comparison behind the “rival Chinese models” headline, and it is genuinely close. Reflection Beam leads on 8 of 13 shared rows. On DeepSWE v1.1 it scores 44.4 to GLM 5.2’s 44.0. It trails on Terminal Bench v2.1 (80.1 to 81.0), AIME 2026, Humanity’s Last Exam, CriPT and GPQA Diamond.

Against the newer Chinese models

Against GLM 5.3, Qwen 3.8 Max and Kimi K3, Reflection Beam is behind on almost everything. It wins none of the 12 shared rows with GLM 5.3 and none of the 13 with Qwen 3.8 Max. Reflection says as much in its own post: “frontier open models like Kimi K3 remain ahead on raw capability”.

Two rows that show the gap

On DeepSWE v1.1, an agentic coding test, Beam scores 44.4 against 61.0 for GLM 5.3, 68.0 for Kimi K3 and 74.2 for DeepSeek V4.1 Flash. On Terminal Bench v2.1, Beam’s 80.1 sits 10.5 points behind DeepSeek V4.1 Flash’s 90.6. Reflection’s pitch is not that Beam wins these races. It is that Beam gets close for less.

What Reflection Beam Does Beyond Coding

reflection beam open weight model chinese rivals d little free library box on a post door open

Coding gets the headlines, but most enterprise agent work is tool calling, search and reading long documents. Reflection’s table covers those too.

Tool calling, search and long context

The table compares Reflection Beam with the best score any open model posts on each row. The gap column is our subtraction.

BenchmarkReflection BeamBest open score in the tableGap
MCP Atlas (tool use)78.784.5 (Qwen 3.8 Max)5.8
tau3 banking (agent tasks)38.055.2 (Qwen 3.8 Max)17.2
AutomationBench public37.054.8 (DeepSeek V4.1 Flash)17.8
BrowseComp with context management77.491.2 (Kimi K3)13.8
AA-LCR (long context)79.388.7 (Kimi K3)9.4
LongBench v265.566.3 (Qwen 3.8 Max)0.8
IFBench (instruction following)79.782.8 (Qwen 3.8 Max)3.1

On long documents Reflection Beam is close to the leaders: 0.8 points behind on LongBench v2 and 3.1 behind on instruction following. The weak spots are agentic task benchmarks such as tau3 banking and AutomationBench, where it trails by more than 17 points. For a model sold as an agent “workhorse”, those are the rows buyers should test first.

The demos

Reflection showed Beam building a live New York subway map from public transit data, writing a 3D browser game, and preparing a fine-tuning notebook for Google’s small Gemma-4 model that raised its accuracy on a held-out text-to-SQL test by 66.5%. On a viral land-and-water mapping puzzle, Reflection says Beam got 95.5% coverage right, between two Anthropic models it cites at 92.5% and 97.8%. Demos are chosen by the vendor, so treat them as illustrations rather than evidence.

Data quality as the hidden lever

Reflection says its curation pipeline removes about 95% of raw internet tokens, yet keeps roughly 1.8 trillion high-quality tokens that conventional filters would have discarded, including 87% of its curated web code. If those claims hold up in the technical report, curation may explain more of Reflection Beam’s efficiency than the architecture does.

The Compute Claim Behind Reflection Beam

reflection beam open weight model chinese rivals e stirling engine running on a coffee mug

Reflection’s central selling point is efficiency: scores “comparable to GLM-5.2 while using 3-4x less inference compute” on advanced reasoning benchmarks. It is worth understanding exactly what that number measures.

How the estimate is built

Reflection estimates generation compute as 2 x active parameters x mean generated tokens per attempt, counting reasoning tokens and the final answer. Scores and token counts for other models come from Artificial Analysis and DataCurve. Reflection calls the result “an approximate compute comparison rather than measured inference cost”.

Doing the arithmetic

Per generated token, Reflection Beam needs about 46 billion operations (2 x 23 billion) and GLM 5.2 about 80 billion (2 x 40 billion). That is a 1.74x advantage from the smaller active size alone (80 divided by 46). To reach 3x to 4x overall, Beam must also produce roughly 1.7 to 2.3 times fewer tokens per answer (3 divided by 1.74, and 4 divided by 1.74). In other words, more than half the claimed saving comes from shorter reasoning, not from the architecture.

Estimated operations per generated token (2 x active parameters), billions
GLM 5.2, 40B active 80
Reflection Beam, 23B active 46

What the estimate leaves out

The formula excludes prompt processing (prefill), attention costs that grow with context length, and serving overhead. For long-context agent work, where a 1 million token window is the selling point, those excluded costs can be large. Reflection trained a controllable length penalty and exposes a reasoning effort setting, so real token use will also depend on how you configure it.

What it means for your bill

Treat the 3-4x figure as a hypothesis to test, not a price. If you run Reflection Beam on your own hardware, memory (driven by the 501 billion total parameters) may cost you more than compute. If you buy it through a hosting partner, the per-token price will decide the economics, and Reflection has not published one.

How Reflection Trained Beam

reflection beam open weight model chinese rivals f shipyard gantry crane over a hull on the slipway

The blog post is unusually open about training. The headline is that Reflection spent more GPU time on reinforcement learning than on pretraining.

Pretraining versus RL

Beam was pretrained “in under four weeks” on 6,144 Nvidia GB300 GPUs, at most 24,576 GPU-weeks (6,144 x 4). The RL phase then ran on 10,500 GB300s for four weeks, or 42,000 GPU-weeks (10,500 x 4). The RL stage therefore used at least 1.7 times the GPU time of pretraining (42,000 divided by 24,576).

Reflection Beam training compute, GPU-weeks (GPUs x weeks, from Reflection’s figures)
Reinforcement learning phase, 10,500 GPUs x 4 weeks 42,000
Pretraining, 6,144 GPUs x under 4 weeks (upper bound) 24,576

The scale of the RL run

The RL campaign generated more than 100 million rollouts, used about 1.3 billion sandboxes for training and grading, and drew on nearly one million coding, agentic and STEM environments. For comparison, Reflection says Inkling was trained on 30 million rollouts and Xiaomi’s MiMo on 753,000. Reflection says capabilities “continued to improve” with no sign of a plateau.

Infrastructure details worth noting

The team sustained an average of 110,000 concurrent rollouts and up to 170,000 concurrent sandboxes. New weights reached the inference fleet in a median of about 12 seconds, 71 inference incidents were absorbed without stopping training, and pretraining reached 92.3% goodput. Those are the operational numbers of a lab that intends to ship more models, and Reflection says the next one is already training.

Behaviour that generalised

Reflection reports that Beam’s browsing improved even though browsing tasks were not in that RL mix. With web access, it says, the model learned on its own to look up and question other AI chatbots, and to call OCR services to read documents. That is a useful agent skill. It is also a cybersecurity consideration: a model that reaches out to other services on its own needs network controls and logging.

Why Reflection Beam Is Aimed at Chinese Open Models

The best open-weight models of the past two years have mostly come from Chinese labs, including DeepSeek, Alibaba’s Qwen, Z.ai and Moonshot’s Kimi. Reflection’s founding pitch was to give the US and its allies a frontier-class open alternative.

The funding behind the pitch

Reflection raised $2 billion at an $8 billion valuation in October 2025, with Nvidia among the investors, and has since signed compute deals with SpaceX and Nebius that TechCrunch puts at more than $7 billion combined, securing Nvidia GB300 chips through 2029. Reflection Beam is the first product that money has produced.

AI factories and sovereign customers

Reflection is targeting enterprises and governments with “AI factories”: systems that let an institution train Reflection’s models on its own data and run them locally. It is already testing the idea with South Korea’s Shinsegae Group. For sovereign buyers, an Apache 2.0 model from a US lab answers a procurement question that a Chinese model sometimes cannot, regardless of benchmark scores.

The open-weight debate

We looked at how founders weigh open against closed models in our report from TechCrunch Disrupt’s open versus closed AI debate, and at the safety questions open weights raise in our piece on open-weight AI safety. Reflection Beam lands in the middle of both arguments.

Running Reflection Beam Yourself: The Hardware Arithmetic

Open weights mean you can download and host the model. A mixture-of-experts design keeps compute per token low, but every expert still has to sit in memory.

Memory for the weights alone

The table multiplies 501 billion parameters by the bytes each number takes at common precisions. It covers weights only, before the memory needed for context (the KV cache), which grows with long prompts.

PrecisionBytes per parameterWeights in memory80 GB accelerators needed for weights
BF162about 1,002 GB13 (1,002 divided by 80 is 12.5)
FP81about 501 GB7 (501 divided by 80 is 6.3)
4-bit0.5about 251 GB4 (251 divided by 80 is 3.1)

What that means in practice

Even at FP8, Reflection Beam needs a multi-GPU server to hold its weights, and in practice a full eight-GPU node once context is added. That is data-centre hardware, not a workstation. Our local LLM hardware guide covers what smaller teams can realistically run.

Hosted options

Reflection says it will launch with “an ecosystem of distribution partners” across hyperscalers and neoclouds, plus integrations with open-source libraries. For most UK businesses, a hosted endpoint will be the sensible way to test Reflection Beam before deciding whether self-hosting is worth it.

What Is Still Missing From the Reflection Beam Launch

The announcement is a preview. Several things that buyers and researchers need do not exist yet.

Weights, report and model card

Reflection says the weights, technical report, model card and developer artifacts will arrive “later this month”. Until then, nobody outside the waitlist can run Reflection Beam, and nobody can reproduce the benchmark table.

Independent verification

TechCrunch noted that Reflection’s performance claims “haven’t been independently verified”. The competitor scores come from third-party sources, while Beam’s own scores come from Reflection. That mix is common at launch, but it is a reason to wait for independent runs.

Safety results

Beam is “undergoing final red-teaming and evaluations”. Reflection trained a separate safety and alignment model from the same checkpoint, merged it with the main model through distillation, and promises to publish the results with the technical report and open-source its internal safety evaluations.

What Reflection Beam Means for UK Businesses

For a UK organisation, Reflection Beam is interesting less for its scores than for what it represents: a credible, permissively licensed, non-Chinese open model at the 500 billion parameter scale.

When it is worth a look

If you need to keep data on your own infrastructure or in a UK region, or your procurement rules exclude Chinese-origin models, Reflection Beam is worth putting on the shortlist once the weights ship. Apache 2.0 allows commercial use and modification.

How to evaluate it

Test it on your own tasks, at the reasoning effort you would actually use, and measure tokens per answer and latency, not just accuracy. Compare it with a closed API and with at least one Chinese open model on the same prompts. Our AI strategy team runs that kind of side-by-side evaluation.

Controls to put in place

Log every outbound call an agent built on Reflection Beam makes, given its learned habit of querying other services. Restrict network access by default and review the model card’s safety findings before any customer-facing use.

Reflection Beam: Frequently Asked Questions

What is Reflection Beam?

Reflection Beam is Reflection AI’s first open-weight model: a text-only mixture-of-experts model with 501 billion total parameters, 23 billion active per token and a 1 million token context window, built for coding and agentic tasks.

Is Reflection Beam better than Chinese open models?

It depends which ones. In Reflection’s own table, it beats GLM 5.2 on 8 of 13 shared benchmarks but loses every shared benchmark to GLM 5.3 and Qwen 3.8 Max, and 13 of 14 to Kimi K3.

Can I download Reflection Beam now?

Not yet. Reflection says the weights will be released under Apache 2.0 later in October 2026. For now there is a waitlist.

How much compute does Reflection Beam save?

Reflection estimates 3-4x less inference compute than GLM 5.2 on advanced reasoning tasks. About 1.74x of that comes from fewer active parameters; the rest depends on generating fewer tokens.

What hardware does Reflection Beam need?

Roughly 501 GB for the weights at FP8 or about 1 TB at BF16, before context memory. That means a multi-GPU server or a hosted endpoint.

References