Byteification turns Qwen3 and Llama 3 into byte-level models

The tokenizer is the last piece of the language-model pipeline that was never learned end to end — and a Nature paper out this week shows it can be swapped out for less than one percent of a pretraining budget.
Researchers have retrofitted four open subword language models into true byte-level models using less than 1% of a typical pretraining budget, and the result — published 7 October in Nature — closes the performance gap that has kept byte-level models a niche curiosity for nearly a decade. Benjamin Minixhofer and colleagues from the Allen Institute for AI, the University of Cambridge, the University of Washington, Imperial College London and LMU Munich introduced "byteification": a two-stage conversion that turns Olmo 3 7B into Bolmo 7B, Qwen3 8B into Bwen 8B, Llama 3 8B into Blama 8B, and OLMo 2 1B into Bolmo 1B by spending just 49.1 billion tokens of extra training. The first stage trains new byte-level components to exactly mimic the frozen original model; the second unfreezes everything and teaches the system to actually use character-level detail.
The numbers justify the effort. Bolmo 7B posts a 16.5-point absolute gain on STEM tasks over BLT 7B, a byte-level model trained from random initialization, while staying close to its subword parent on general benchmarks. On character-level understanding — the weakness that makes subword models clumsy with code, biological sequences and anything where meaning lives in individual characters — the byteified models don't just match their sources, they beat them, and they beat every earlier public byte-level model by wide margins. Bwen 8B, converted from Qwen3, performs close to and sometimes surpasses the original.
What makes this more than a paper artifact is that the converted models inherit the entire ecosystem of the model they came from. Merging a post-trained Olmo 3 instruction checkpoint into Bolmo through task arithmetic — plain weight subtraction and addition, no training at all — lifted Bolmo's instruction-following to parity with the original checkpoint. Byte-level models also sidestep the softmax bottleneck: the team showed a subword model becomes uneconomical once its vocabulary passes roughly 200,000 to 400,000 entries, while a byteified model simply patches more bytes per step and keeps getting faster.
We covered Meta's earlier from-scratch attempt — Meta's distilled byte models break through the token ceiling — and byteification is the cheaper road to the same destination: rather than pretraining a byte model to catch up, convert a model you already trust. All checkpoints, including the intermediate stage-one weights, are released openly.
What to watch: the method is proven at 1B and 7B parameters — the real test is whether someone applies it to a frontier model with a full post-training pipeline behind it.
If the tokenizer was the last hand-fixed component of an LLM, what else in the training stack is still a hand-me-down from 2018? Tell us in the comments.



