Writing
Hi :) this is where I overshare about AI.
Lab deep dive · Ai2
The model weights are open. Can we see how the model learned?
Why open model training matters, and how to test the steps that shape an assistant.
Open models

Paper deep dive
A million tokens can still lose the right document
Why seeing every document is not the same as finding the right one.
Retrieval

Paper deep dive
The agent did everything right. The website still said no.
What a verifier can learn when the assistant took the right steps but the website failed.
AI Agents
Paper deep dive
Can one reward teach an assistant everything?
A reading of MAHALO, where accuracy, human values and tutoring pull an assistant in different directions.
AI Alignment

Reading notes
The one percent we stopped looking at
Muon, Never Give Up, and the cases an average can hide.
Machine Learning

Paper deep dive
Never Give Up: why easy problems eat the RL budget
A deep dive into the Matthew Effect, and how adaptive sampling gives hard problems another chance.
Reinforcement Learning
Paper deep dive
A coin flip made this model better at math. Or did it?
Notes on "Spurious Rewards", a paper about what RL training really teaches a model.
Reinforcement Learning

My research
The polite sentence that broke my model
How we taught small language models to call tools with GRPO, and why the strictest rule worked best.
Reinforcement Learning

Coming next
Paper deep diveOne research paper, explained simply
Lab mapWhat one AI lab is working on, and why it matters