
The Frontier
Two fixed opening tokens push a base model past its RL version
A new paper from MIT, UC Berkeley, Washington University and the Allen Institute for AI says much of what reinforcement learning teaches a model may come down to how it starts an answer. One good prefix, fixed in advance, is enough to close the gap. Fix the first two tokens and Olmo-3-7B beats its own RL-trained twin. The paper — "Base Models Can Reason By Taking a Cue From Training Data" — tested what happens when researchers pre-fill a model's response opening instead of letting it choose. O










