Two fixed opening tokens push a base model past its RL version

Share
Two fixed opening tokens push a base model past its RL version

A new paper from MIT, UC Berkeley, Washington University and the Allen Institute for AI says much of what reinforcement learning teaches a model may come down to how it starts an answer. One good prefix, fixed in advance, is enough to close the gap.

Fix the first two tokens and Olmo-3-7B beats its own RL-trained twin. The paper — "Base Models Can Reason By Taking a Cue From Training Data" — tested what happens when researchers pre-fill a model's response opening instead of letting it choose. On MATH-500, the base Olmo-3-7B scores about 42% on its own; force the opening to . followed by a blank line and "Okay," and it jumps to roughly 78%, ahead of the reinforcement-learning version's 75%. Qwen3-14B shows the same pattern with " Alright," as the cue, climbing from 72% to 87% and matching its RL counterpart. No parameters were touched and no worked examples were added — only the starting words changed. The researchers picked the cues by scanning openings the model already produces and measuring which one yields the steadiest answers; on correct responses, about half begin with a period and two newlines, versus 14% of wrong ones.

It reframes what RL is actually buying you. Under an RL-Zero setup, the biggest distribution shift between the base and RL-trained checkpoints sits in the first two tokens of the response: RL raises the odds of the winning cue from 0.14 to 0.65 for Olmo and from 0.04 to 0.58 for Qwen. A model given the RL version's opening, then left to continue on its own, matches the RL score — and a fixed cue is worth roughly 100 steps of RL training for Olmo. The authors' data-rewiring experiment makes the mechanism concrete: swap "Okay" for "Chicken" in the mid-training corpus, retrain, and .\n\nChicken becomes an effective reasoning cue too, lifting MATH-500 from 18% to 77%. "Think step by step" works the same way after "duck duck goose" is substituted. The cue is not magic wording — it is a pointer into associations the training data built.

The safety finding is the part worth watching. Openings also steer refusal behavior: an "I'm sorry" start makes Olmo refuse more, including harmless requests, while "Okay," loosens compliance on dangerous ones — and merely changing punctuation between Okay, and .\n\nOkay flips the tendency. That means alignment behavior can hinge on a prefix a user may be able to influence, not just on learned weights. The caveat is scope: the results cover two open models on math and code benchmarks, so generalization to frontier models is unproven.

What to watch: whether labs test prefix sensitivity as an eval, and whether the cue trick survives at frontier scale.

Does RL mostly teach models how to start talking? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t