Series

30 days, 30 papers.

One paper a day, explained simply.

Day 1: Stop agents from gaming the test
Day 1

Stop agents from gaming the test

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Day 2: Teach models to win clean
Day 2

Teach models to win clean

MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement

Day 3: A week old, already tested
Day 3

A week old, already tested

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Day 4: The rich get richer
Day 4

The rich get richer

Never Give Up: Learning to Solve Hard Problems in RL for LLMs

Day 5: Grade the process, not the output
Day 5

Grade the process, not the output

OSWorld-Pro: evaluating where computer-use agents fail

Day 6: Train the search, keep the model frozen
Day 6

Train the search, keep the model frozen

Beyond Repeated Sampling: concepts guide the search