Writing

Hi :) this is where I overshare about AI.

Lab deep dive · Ai2

The model weights are open. Can we see how the model learned?

Why open model training matters, and how to test the steps that shape an assistant.

8 min readOpen models
Open training notebook
Paper deep dive

A million tokens can still lose the right document

Why seeing every document is not the same as finding the right one.

10 min readRetrieval
Wall of documents
Paper deep dive

The agent did everything right. The website still said no.

What a verifier can learn when the assistant took the right steps but the website failed.

8 min readAI Agents
Train ticket attempt: steps completed, verification failed, ticket not booked
Paper deep dive

Can one reward teach an assistant everything?

A reading of MAHALO, where accuracy, human values and tutoring pull an assistant in different directions.

7 min readAI Alignment
Can one reward teach an assistant everything? cream and teal cover
Reading notes

The one percent we stopped looking at

Muon, Never Give Up, and the cases an average can hide.

7 min readMachine Learning
The one percent we stopped looking at cover
Paper deep dive

Never Give Up: why easy problems eat the RL budget

A deep dive into the Matthew Effect, and how adaptive sampling gives hard problems another chance.

8 min readReinforcement Learning
Never Give UpFour attempts are enough for the easy problem. The hard problem stays in the queue and gets another chance.PAPER DEEP DIVE · NEVER GIVE UPStop letting the easy oneseat the budget.EASY PROBLEM✓ ✓ ✓ ✓→ stopHARD PROBLEM✗ ✗ ✗ ✗→ try againdhruvi paprunia · notes on RL
Paper deep dive

A coin flip made this model better at math. Or did it?

Notes on "Spurious Rewards", a paper about what RL training really teaches a model.

3 min readReinforcement Learning
My research

The polite sentence that broke my model

How we taught small language models to call tools with GRPO, and why the strictest rule worked best.

6 min readReinforcement Learning
Coming next
Paper deep diveOne research paper, explained simply
Lab mapWhat one AI lab is working on, and why it matters