Byte-level model research reached the pages of Nature on 7 October 2026, and the headline result is one most people can test in a chat window: ask a large language model to spell a word backward and it often fumbles. A team from LMU Munich, the Allen Institute for AI (Ai2), the University of Cambridge, the University of Washington and Imperial College London has shown a cheap way to fix that. They convert an existing model so that it reads raw bytes, roughly one per letter, instead of the word fragments it was built on.

Tech Xplore carried the university’s announcement on 9 October under the headline above. The underlying paper, “Retrofitting language models to operate over bytes”, is open access, and it is more interesting than a spelling trick. It claims to remove “a long-standing performance barrier” that has kept every byte-level model behind conventional models for years.

This article explains why AI models struggle with letters, what the researchers built, how the new models score against their originals, what they give up, and what any of this means for businesses that rely on language models for code, data and documents.

Why a Byte-Level Model Matters: AI Models Cannot See Letters

byte-level model - byte level model spelling backward ai sees every letter b flour sifter over a mixing bowl

Before a language model reads anything, a component called a tokenizer chops the text into chunks. Most chunks are whole words or pieces of words, taken from a fixed vocabulary that the paper puts at between 30,000 and 300,000 entries. The model then works on numbered chunks, not on letters. A byte-level model skips that step and reads the bytes that computers use to store text.

What subword tokenization does to a word

The LMU announcement gives a neat example: the model behind ChatGPT splits “LMU München” into three chunks, “LM”, “U” and “München”, so it “never directly sees the individual letters in ‘München'”. We checked that claim with OpenAI’s open-source tiktoken library, using the o200k_base encoding that OpenAI introduced with GPT-4o. It reproduces the split exactly.

TextUTF-8 bytesTokensHow o200k_base splits it
strawberry103st · raw · berry
LMU München123LM · U · ␣München
München (no leading space)83M · ün · chen
spelling82sp · elling
byteification132byte · ification
ATCGGATCCA (a DNA fragment)105AT · CG · G · AT · CCA
def reverse(s):154def · ␣reverse · (s · ):

The table counts bytes with Python’s UTF-8 encoder and tokens with tiktoken 0.14. Note the two rows for München. With a space in front, the word is one token; without it, the same word is three different tokens. The paper describes a related problem it calls tokenization bias, where prompts that end mid-word or with a space behave unexpectedly, and it is one reason the authors want models that see the raw text.

Why spelling backward trips up AI models

To reverse “strawberry”, a subword model has to recall which letters sit inside “st”, “raw” and “berry”, then emit them in the opposite order. It never saw those letters as separate inputs. It learned them indirectly, from patterns across billions of documents. That works for common words and breaks down for rare ones, names and invented strings.

Research cited in the paper backs this up. A 2025 study titled “The strawberry problem” found that character-level understanding emerges mainly through scale. The CUTE benchmark, built by Lukas Edman, Helmut Schmid and Alexander Fraser, measures how well models understand the characters inside their tokens, and subword models consistently fall short on it.

The problems beyond spelling

The Nature paper lists four costs of subword tokenization. Character information is lost. Prompts that end mid-word or with a space can produce odd output. A fixed vocabulary favours English. And every token gets the same compute, however much information it carries.

A byte-level model can, in principle, address all four. The first problem matters most for scientific data. As the authors put it, meaning in “computer code or biological sequences” depends “on the individual characters or bytes”.

What the Nature Study Found About Byte-Level Model Training

byte level model spelling backward ai sees every letter c rear view mirror on a windscreen mount

Byte-level models are not new. Google’s ByT5, MambaByte, SpaceByte and Meta’s Byte Latent Transformer all read bytes, and several claimed better efficiency than subword models. We covered the architecture in our guide to the Byte Latent Transformer. Yet, as the paper notes, “all leading LLMs still exclusively rely on subword tokenization”.

The byte-level model gap the authors set out to close

The authors’ diagnosis is about money and pace, not architecture. Earlier work trained each byte-level model from scratch and compared it with a subword model also trained from scratch. Meanwhile, the best subword models improve fast, through better training data, architecture and post-training. “Keeping up with this pace is infeasible for byte-level LLM development without extensive investments,” they write.

Their answer is to stop starting from scratch. Instead of building a byte-level model from random weights, they take a finished subword model and retrofit it.

Byteification in one sentence

The authors call the process byteification: take an existing subword model, keep its expensive core, and teach new byte-handling components to feed it. The whole conversion used 49.1 billion tokens of training, which the paper describes as “less than 1% of a typical pretraining budget”.

That claim checks out against the source model. Ai2’s model card says Olmo 3 7B was pretrained on 5.93 trillion tokens. Dividing 49.1 billion by 5.93 trillion gives 0.83%. Hofmann, the study’s last author and a junior professor at LMU Munich, summed up the aim in the university’s release: “We believe that byteification, and byte-level LLMs more generally, have the potential to overcome some long-standing shortcomings of LLMs.”

Four byteified models

The team converted four models. Each byte-level model keeps its parent’s core and ends up within about 5% of its parameter count, because the large subword output table is removed while new byte-handling layers are added.

Byte-level modelConverted fromParameters (byte vs source)Change
Bolmo 7BOlmo 3 7B (Ai2)7.63B vs 7.30B+4.5%
Bolmo 1BOLMo 2 1B (Ai2)about 10M fewer−0.7%
Bwen 8BQwen3 8B Base (Alibaba)8.31B vs 8.19B+1.5%
Blama 8BLlama 3 8B (Meta)8.25B vs 8.03B+2.7%

Parameter counts come from Table 1 of the paper and include embeddings; the percentage changes are the paper’s own. Bolmo 7B and Bolmo 1B were first released by Ai2 in December 2025 alongside a preprint. Bwen and Blama are new in the Nature version, and they matter. They show the method is not tied to one lab’s models.

How Byteification Builds a Byte-Level Model From an Existing LLM

byte level model spelling backward ai sees every letter d ice cube tray with cubes popped out

The design follows a family the authors call latent tokenizer language models. The model still groups text into chunks, which the paper calls patches, but it learns where the chunk boundaries go, and it can always look inside them.

The parts of a byteified byte-level model

Text flows through five stages:

  • Byte embeddings. Each byte gets a vector. The model also adds the embedding of the longest matching subword from the original vocabulary, so the old knowledge is not thrown away.
  • Local encoder. One lightweight mLSTM layer, a fast recurrent design, adds context to each byte.
  • Boundary predictor. It decides where each patch ends, using one byte of future context.
  • Global model. The patches pass through the original transformer, which holds most of the compute and knowledge.
  • Local decoder. Four mLSTM layers turn patch outputs back into byte predictions, including where the next boundary falls.

One byte of lookahead

The cleverest change is small. Earlier models of this type decided boundaries using only the text seen so far. The authors point out that ordinary tokenizers cheat: they look ahead. Given “Hello Wor!”, a tokenizer ends a chunk after “Wor”; given “Hello World!”, it does not, even though the text is identical up to that point.

So the new boundary predictor peeks one byte ahead while reading a prompt. During generation, where the next byte is unknown, the decoder predicts boundaries itself. The authors found that one byte of lookahead was “largely sufficient” to match subword behaviour. After stage 1, the predictor matched the original tokenizer’s boundaries with over 99% accuracy.

Two training stages

StageWhat trainsDataGoal
1New byte components only; original transformer frozen9.8B tokens (about 43B bytes)Exactly mimic the source model
2Whole network39.3B tokens (about 173B bytes)Learn to use byte-level information
Total49.1B tokens0.83% of Olmo 3 7B’s 5.93T pretraining tokens

The token and byte figures are the paper’s; the total and the percentage are our arithmetic. The same figures show how much longer byte sequences are: 173 billion bytes divided by 39.3 billion tokens is 4.40 bytes per token, and stage 1 gives 4.39. Stage 1 is cheap because it only back-propagates through the first four layers of the frozen transformer. The paper found the smaller 1B model gained more from stage 1 than the 7B one.

The spelling lessons in the training data

There is a detail here that matters for the headline. The training mix of about 172 billion tokens from Ai2’s Dolma 3 data was “augmented with 75M tokens of CUTE-style data”, roughly 0.04% of the total. Those synthetic exercises included “spelling out words, reversing words as well as swapping, deleting and substituting characters”. They were generated from a list of 150,000 words, with “zero overlap” with the benchmark’s test words.

In other words, the byte-level model was explicitly taught to spell backward. As the next section shows, that teaching explains a large share of the improvement.

Spelling Backward Benchmarks: The Byte-Level Model Scorecard

byte level model spelling backward ai sees every letter e bead loom with a half woven band

The paper measures character skills with two benchmarks. CUTE tests character understanding in English. EXECUTE, by the same group, extends that to many other languages. The paper does not publish a separate score for reversal alone, so these two are the closest public measure of the “spelling backward” claim.

CUTE scores, byte-level model vs original

CUTE character-understanding score, 7B-8B models (Nature, Table 1)
Blama 8B (byteified Llama 3) 87.9
Bwen 8B (byteified Qwen3) 87.2
Bolmo 7B (byteified Olmo 3) 80.9
Olmo 3 7B with the same extra training (control) 74.6
Qwen3 8B (original) 66.9
Llama 3 8B (original) 58.3
Olmo 3 7B (original) 54.7
BLT 7B (earlier byte model, Meta) 48.0

Bar widths equal the score out of 100. The gains over each parent are 29.6 points for Blama (87.9 − 58.3), 26.2 for Bolmo (80.9 − 54.7) and 20.3 for Bwen (87.2 − 66.9). Ai2’s December blog, written about the preprint, put Bolmo’s gain on its wider character aggregate at “nearly twenty points”; in the Nature table that aggregate reads 78.1 against 54.3, a gap of 23.8.

The control model tells the real story

The most useful number in the chart is the control. The authors trained the original Olmo 3 on exactly the same data mix, character exercises included, without converting it. That subword model jumped from 54.7 to 74.6 on CUTE, a gain of 19.9 points. The byte-level model added a further 6.3 points on top (80.9 − 74.6).

So most of Bolmo’s improvement over Olmo 3 came from the spelling exercises, and the smaller remainder came from the byte architecture. That does not make the result less real, because the authors also report that “byte-level models otherwise do not acquire this knowledge through our short training schedule”. It does mean the honest summary is “training data plus bytes”, not “bytes alone”.

Gains in other languages

EXECUTE results are more striking, because the character exercises were English only. Bolmo 7B scored 75.4 against 53.8 for Olmo 3, and Blama 74.1 against 56.2 for Llama 3. The control Olmo 3 reached 66.9, so here the byte architecture added 8.5 points over the same data. The authors read this as evidence that English practice “can help acquire generalizable knowledge about the characters within words”.

Earlier byte models did not win on letters

One finding cuts against intuition. The earlier byte-level models, EvaByte 6.5B, TFree-HAT 7B and BLT 7B, scored 47.7, 54.5 and 48.0 on CUTE. None beat the original subword Olmo 3 at 54.7. Reading bytes is not enough on its own; the paper suggests character skill is “acquired primarily through scale”, which a heavily trained subword model can buy with sheer volume.

What a Byte-Level Model Gives Up: Code, Maths and Speed

byte level model spelling backward ai sees every letter f laced jacquard punch cards in a zigzag stack

The researchers do not claim a free lunch, and their table shows the costs plainly. Across the full 7B evaluation suite, the averages are almost identical: Bolmo 7B scored 62.5 against 62.3 for Olmo 3, with a bootstrap P value of about 0.44, meaning no significant difference. Bwen 8B scored 69.8 against 70.3 for Qwen3, and Blama 59.8 against 60.7.

Where the scores slipped

Benchmark groupBolmo 7BOlmo 3 7BBwen 8BQwen3 8B
Whole suite (average)62.562.369.870.3
Character understanding78.154.378.966.2
Code, first attempt (pass@1)27.631.139.645.8
Code, best of 16 (pass@16)56.449.769.060.7
HumanEval pass@140.649.062.971.0
Maths48.955.360.268.2
GSM8K word problems68.073.176.784.7
Multiple choice, STEM65.566.377.078.9

All figures are from Table 1 of the Nature paper. Maths fell by 6.4 points for Bolmo (55.3 − 48.9) and 8.0 for Bwen (68.2 − 60.2). HumanEval, a standard coding test, fell by 8.4 and 8.1 points. Those are not rounding errors, and anyone choosing a byte-level model for maths or coding should weigh them.

The pass@16 puzzle

The code results point in two directions. On the first attempt, the byte-level model is worse. Given 16 attempts, it is better: Bolmo scored 56.4 against 49.7 and Bwen 69.0 against 60.7. The authors read this as more diverse output, but they warn that proving it “would require mapping the quality-diversity Pareto frontier”, because sampling settings may affect byte models differently. Treat it as a lead, not a finding.

Speed

Byte sequences are about 4.4 times longer than token sequences, so speed is the obvious worry. The preprint version of the paper measured Bolmo decoding at about 125 bytes per second against about 150 for the subword model at the same compression, on one H100 GPU. Reading a 72,000-byte prompt took about 1 second against 0.8 seconds. That is roughly 17% slower generation (25 ÷ 150) and 25% slower reading (0.2 ÷ 0.8): close, not equal.

Why a Byte-Level Model Could End Up Faster

The paper’s most commercially interesting argument is about the future. A subword model gets faster by using bigger tokens, which means a bigger vocabulary. But every extra vocabulary entry makes the final softmax layer, the step that scores every possible next token, more expensive.

The vocabulary ceiling

The authors tested this by stretching OLMo 2 1B’s vocabulary with SuperBPE, a method that adds multi-word tokens, to 200,000 and then 400,000 entries. Somewhere between those sizes, the subword model became, in the paper’s words, “Pareto-dominated”: slower for the same quality. A byte-level model has no such ceiling. Its output layer covers just 512 symbols, the 256 possible byte values plus a version of each marking a patch boundary.

Compression as a dial

Because the byte-level model learns its own patch boundaries, it can be trained to use longer patches on average. The preprint estimates that Bolmo overtakes the subword model in inference efficiency at about 6.6 bytes per patch, against the 4.4 bytes per token of the current setup. Ai2 describes compression as “a toggleable knob”. Interestingly, simple byte-pair-encoding merges beat entropy-based boundaries for this, the opposite of earlier byte-model research.

Reusing the Ecosystem: Post-Training a Byte-Level Model for Free

Building a model is only half the job. Instruction-following, safety tuning and domain fine-tunes come afterwards, and they are expensive. A converted byte-level model would be far less useful if all that work had to be redone in byte space.

The task arithmetic result

The authors tried a shortcut. They took an Olmo 3 checkpoint that had been post-trained with reinforcement learning to follow instructions, subtracted the base model’s weights to isolate what post-training changed, and added that difference to Bolmo’s transformer layers. No extra training was involved.

On the IFEval instruction-following benchmark, Ai2 reports that base Bolmo scored 31.1% against 35.4% for base Olmo 3. After the merge, Bolmo scored 67.4%, against 66.9% for the post-trained Olmo 3. A gain of 36.3 points (67.4 − 31.1) for no training compute is the result most likely to interest companies that maintain fine-tuned models.

IFEval instruction-following accuracy, before and after merging (Ai2)
Bolmo 7B after zero-cost merge 67.4%
Olmo 3 7B, post-trained with RL 66.9%
Olmo 3 7B base 35.4%
Bolmo 7B base 31.1%

Bar widths equal the accuracy percentage. The merged byte-level model ends 0.5 points above the subword checkpoint it borrowed from (67.4 − 66.9), which is within the noise of a single benchmark run, so the fair reading is “level”, not “better”.

The caveat

It does not always work. The trick needs the post-trained model’s embeddings to be “resettable” to the base model’s values “without a large performance loss”, which the paper says is “generally, but not always, the case”. Ai2 calls embedding-friendly post-training “an important area for future work”.

Where a Byte-Level Model Helps First

Hofmann stressed in the LMU release that character skills are “more than an academic curiosity”: “A good representation of the low-level structure of text is critical in many areas of science, for example when working with code or biological sequences.” The tokenizer table earlier shows why. The DNA string ATCGGATCCA became five uneven chunks, none of which line up with the biology.

Code, identifiers and structured data

Software is full of strings where single characters matter: variable names, file paths, hashes, version numbers and regular expressions. So is business data, from part numbers and postcodes to invoice references and IBANs. A model that cannot see individual characters is more likely to transpose or drop one. A byte-level model at least sees what it is editing.

Languages other than English

The paper notes that fixed vocabularies lead “in practice” to “English-centricity”. Our tokenizer test shows the effect in miniature: the eight-letter Greek word Ελληνικά takes 16 bytes and three tokens, while the eight-letter English word “spelling” takes eight bytes and two tokens. A byte-level model treats every language’s bytes the same way, although non-Latin scripts still use more bytes per character in UTF-8.

Tasks that are really about spelling

Spell-checking, name matching, anagram and word-game features, transliteration and accessibility tools all depend on characters. These are the jobs where today’s chat assistants most visibly fail, and where the byteified models show the clearest gains.

What a Byte-Level Model Means for Businesses Using AI Today

None of this changes the AI tools most businesses use this month. Every leading commercial model still uses subword tokens, and the largest model in the study has about 8 billion parameters, far below the frontier. But there are practical lessons now.

Test the character-level tasks in your workflows

If your team uses AI for anything involving exact strings, such as reference numbers, product codes, data cleaning or code review, test that specifically. Do not assume that a model which writes fluent prose can copy an 18-character order number without error. Where accuracy matters, have the model call a tool, such as a short script, rather than manipulate characters itself. Our ML model development team can help design those tests.

Watch the tokenizer in your AI bill

Tokenizers also decide cost, because API pricing is per token. When Anthropic launched Claude Haiku 5.5, its migration guide warned that a new tokenizer produced around 30% more tokens for the same text; we covered it in our Claude Haiku 5.5 launch report. Any future byte-level model sold through an API would need a new pricing unit, and buyers should read it carefully. Our five keys to controlling AI token costs still apply.

Try the open models if you have the use case

Bolmo 7B and Bolmo 1B are on Hugging Face under the Apache 2.0 licence, for “research and educational use” under Ai2’s responsible use guidelines. They run with the standard transformers library (version 4.57.3 or later, with remote code enabled). Hugging Face shows about 1,741 downloads of Bolmo 1B and 589 of Bolmo 7B in the past month. The code is on GitHub as bolmo-core and the training data as bolmo_mix. A 7B model fits on a single workstation GPU; see our local LLM hardware guide.

Limits and Open Questions for the Byte-Level Model

The study is careful, and its limits are worth stating plainly.

Scale and scope

The largest byteified model has about 8 billion parameters. Whether the method holds for frontier-scale models is untested. The character exercises were English only, and the code-diversity result needs more work. The comparison with earlier byte models is also not like for like, since BLT 7B and the others were trained from scratch on smaller budgets.

Benchmarks are not the job

CUTE and EXECUTE are useful, but they are synthetic tests, and benchmark gains do not always survive contact with real work, a theme we explored in the tests that grade AI may be getting it wrong. The training data was designed to resemble CUTE while avoiding its test words, which is legitimate but leaves room for gains that are narrower than the scores suggest.

Will frontier labs follow?

The paper’s strategic point is that byteification makes byte-level research cheap, so promising designs can be found by converting existing models before anyone spends a full pretraining budget. Whether OpenAI, Google, Anthropic or Meta adopt it will depend on the speed and pricing questions above. For now, the byte-level model has gone from “research curiosity”, Ai2’s phrase, to a credible option that anyone can download.

Questions About the Byte-Level Model Study

Why can’t ChatGPT spell words backward reliably?

Models like ChatGPT read text as word fragments, not letters. OpenAI’s o200k_base tokenizer turns “strawberry” into three pieces, so reversing it means recalling letters the model never saw as separate inputs.

What is byteification?

It is the researchers’ name for converting an existing subword model into a byte-level model with a short, two-stage training run. It used 49.1 billion tokens, about 0.83% of Olmo 3 7B’s pretraining.

Is a byte-level model better than a normal LLM?

On character tasks, clearly: Bolmo 7B scored 80.9 on CUTE against 54.7 for Olmo 3. On average across all tasks they are level, with losses in maths and first-try coding.

Which models were converted?

Olmo 3 7B (into Bolmo 7B), OLMo 2 1B (Bolmo 1B), Qwen3 8B (Bwen 8B) and Llama 3 8B (Blama 8B).

Can businesses use Bolmo today?

The models are open under Apache 2.0 and intended for research and education. They suit experiments with code, multilingual text or string-heavy data, not production chat assistants.

References