Byteification turns Qwen3 and Llama 3 into byte-level models

Share
Byteification turns Qwen3 and Llama 3 into byte-level models

The tokenizer is the last piece of the language-model pipeline that was never learned end to end — and a Nature paper out this week shows it can be swapped out for less than one percent of a pretraining budget.

Researchers have retrofitted four open subword language models into true byte-level models using less than 1% of a typical pretraining budget, and the result — published 7 October in Nature — closes the performance gap that has kept byte-level models a niche curiosity for nearly a decade. Benjamin Minixhofer and colleagues from the Allen Institute for AI, the University of Cambridge, the University of Washington, Imperial College London and LMU Munich introduced "byteification": a two-stage conversion that turns Olmo 3 7B into Bolmo 7B, Qwen3 8B into Bwen 8B, Llama 3 8B into Blama 8B, and OLMo 2 1B into Bolmo 1B by spending just 49.1 billion tokens of extra training. The first stage trains new byte-level components to exactly mimic the frozen original model; the second unfreezes everything and teaches the system to actually use character-level detail.

The numbers justify the effort. Bolmo 7B posts a 16.5-point absolute gain on STEM tasks over BLT 7B, a byte-level model trained from random initialization, while staying close to its subword parent on general benchmarks. On character-level understanding — the weakness that makes subword models clumsy with code, biological sequences and anything where meaning lives in individual characters — the byteified models don't just match their sources, they beat them, and they beat every earlier public byte-level model by wide margins. Bwen 8B, converted from Qwen3, performs close to and sometimes surpasses the original.

What makes this more than a paper artifact is that the converted models inherit the entire ecosystem of the model they came from. Merging a post-trained Olmo 3 instruction checkpoint into Bolmo through task arithmetic — plain weight subtraction and addition, no training at all — lifted Bolmo's instruction-following to parity with the original checkpoint. Byte-level models also sidestep the softmax bottleneck: the team showed a subword model becomes uneconomical once its vocabulary passes roughly 200,000 to 400,000 entries, while a byteified model simply patches more bytes per step and keeps getting faster.

We covered Meta's earlier from-scratch attempt — Meta's distilled byte models break through the token ceiling — and byteification is the cheaper road to the same destination: rather than pretraining a byte model to catch up, convert a model you already trust. All checkpoints, including the intermediate stage-one weights, are released openly.

What to watch: the method is proven at 1B and 7B parameters — the real test is whether someone applies it to a frontier model with a full post-training pipeline behind it.

If the tokenizer was the last hand-fixed component of an LLM, what else in the training stack is still a hand-me-down from 2018? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t