← All writing
Reinforcement Learning

A coin flip made this model better at math. Or did it?

Notes on "Spurious Rewards", a paper about what RL training really teaches a model

Dhruvi Paprunia3 min read · Also on Medium ↗
Coin flip reward and math performance illustration

Reinforcement learning is supposed to work because of the reward. You tell the model what a good answer looks like, and it learns to give more of them. Get the reward wrong and the model should learn the wrong thing.

In "Spurious Rewards: Rethinking Training Signals in RLVR", Sewon Min and her co-authors put that idea to the test. They trained a math model with rewards that made no sense. Some were random. Some rewarded only wrong answers. The model still got much better at math.

I spent an evening with this paper, and the reason it works is more interesting than the result itself.

First, what is RLVR?

RLVR stands for reinforcement learning with verifiable rewards. You give a model a question with a checkable answer, like a math problem. The model tries a few answers, a program checks each one, and correct answers earn a reward of 1. Wrong answers get 0.

GRPO is one popular way to do this. It samples a group of answers for the same question and pushes the model toward the ones that scored above the group average.

The idea is simple: the reward tells the model what "good" looks like.

The surprising result

The team took Qwen2.5-Math-7B, a model already trained heavily on math, and swapped the real reward for worse and worse ones:

On the MATH-500 benchmark, the real reward improved accuracy by 29.1 points. The random reward improved it by 21.4 points. A coin flip got most of the way there.

Ground-truth and random reward gains on MATH-500

Why would a coin flip help?

This is my favourite part of the paper. The answer is hiding in one small piece of GRPO called clipping.

Clipping stops the model from changing too much in one step. If a word's probability tries to jump by more than about 20%, the update gets cut off.

But look at what that means for a word the model already likes. Say it picks a word 85% of the time. To be clipped, it would have to go above 85% × 1.2 = 102%. That is impossible, so that word is never held back. Rare words, on the other hand, hit the limit easily.

So even when the rewards are pure noise, clipping quietly favours what the model was already likely to do. The authors show this directly: when they removed clipping, random rewards stopped helping.

Clipping favors high-prior tokens

The habit it turned up

For Qwen2.5-Math, one of its habits happens to be very good. It often "reasons in code", writing Python-style steps even though nothing runs the code. Answers written this way were right 60.9% of the time, compared to 28.0% without.

Before training, 65% of its answers used code reasoning. After training with spurious rewards, it was over 90%. The coin flip didn't teach the model math. It just made the model use its best habit more often.

Code reasoning frequency and accuracy

What this doesn't mean

It doesn't mean a random reward will make any model better. If training turns up a common habit, it can turn up a bad one too. The team tested Llama and OLMo as well; their gains with spurious rewards were mostly flat and sometimes negative. Qwen2.5-Math had a useful code-reasoning habit to begin with. The coin flip didn't create it.

And it doesn't mean the reward method deserves all the credit. Qwen is a common starting point for RLVR research, so a result that looks strong on Qwen may be leaning on the model's existing strengths. That's why the paper tests other model families and calls for future work to do the same.

What I'm taking from this

  1. Check RL results on more than one model family. The paper makes this point strongly.
  2. Use a random reward as a baseline. If random rewards get you most of your gain, your reward isn't the reason your model improved.
  3. Know what the model already does well before training. RL often turns up existing skills more than it teaches new ones.

Having written reward functions myself, I used to think of the reward as the whole story. This paper made me look at the model first.


Read the paper: Spurious Rewards: Rethinking Training Signals in RLVR

Dhruvi Paprunia3 min read · Also on Medium ↗
More writing →