← All writing
Machine LearningReading notes

The one percent we stopped looking at

Muon, Never Give Up, and the cases an average can hide.

Dhruvi Paprunia7 min read
The one percent we stopped looking at. A conceptual 96 percent result highlights four percent worth investigating.

A result can be 96% right and still leave me thinking about the other 4%. Not because every system can be perfect, but because those failures are telling us something. What is different about the cases that are still missed? Did our method make them hard to see, or did we decide they were too expensive to keep trying?

By an edge case, I mean a rare or hard example the system is supposed to handle, not every small direction in a weight update and not every failure someone can invent. My starting point is not to treat those leftover cases as cleanup after training. I would first sort them: some are irreducible noise or outside the task; some are a fixable blind spot. That triage is the work. Only then would I ask whether the reward or training signal gives the important, learnable cases a real chance from the start. If easy cases make up almost all the feedback, a good average can become a reason to stop looking. At scale, even a 1% failure rate would mean 1,000 failed uses in 100,000 attempts. And a person whose use failed does not experience the average.

That is the lens through which I started connecting Muon and Never Give Up. They do not directly assign more reward to the missed cases I want to study. Muon changes the geometry of a weight update; Never Give Up changes which RL problems get more attempts. Neither promises to fix product edge cases. But each asks, in its own way, what the usual training process is underweighting. Comparing them helps me make my own question sharper: what would it take to notice the hard cases while the model is learning, not only after we publish a top-line score?

My approach: find the failure before weighting it

At JioHotstar, working on systems used at streaming scale makes this concrete for me. People do not only test the 80% of cases where a product works; they look for the edge cases that make a user flow fail. If one missed case breaks that flow, the person who hits it does not experience a 99% success rate. They experience a broken product. That is why I want to know which failures are harmless and which ones break trust. A percentage alone cannot tell me.

"Nothing can be perfect" may be true, but it does not tell us why the remaining cases failed, whether the same pattern keeps repeating, or which part is fixable. I would inspect that leftover slice, not assume every miss deserves an identical fix. Once I know what is worth learning, I would test shaping the reward or signal so those cases count during training. But there is a cost: attention and compute given to the 1% are not free. Overweighting it can hurt the average, or make the model chase noise while the cases it already handles get worse. I would measure that tradeoff against the failures it prevents, not assume more weight always helps. If the label is noise or the task is impossible, more reward pressure can teach the wrong thing. The sampling budget matters too. A hard example may need more chances to produce a useful signal before a reward can help it.

A hard example can be missed because it gets too few tries. Muon offers a different lens on what gets drowned out during training: a few directions can dominate a weight update. That is an analogy, not the same failure.

Muon: a comparison about the shape of an update

Training changes a model's weights a little at a time. For a hidden-layer weight matrix, you can think of an update as a set of directions in which those weights might move. Keller Jordan's Muon takes an update made with momentum and approximately orthogonalizes that matrix before applying it. AdamW adjusts each parameter using its gradient history; Muon balances the geometry of the whole matrix update. I think of them as specialists, not a contest with one universal winner. The job determines which tool fits, and even Muon keeps AdamW for parameter types it does not handle.

Muon: dominant and smaller directions in an optimizer update. Rare direction benefit is a hypothesis, not a guarantee.

Here is a picture to build intuition, not the actual Muon math. Imagine several people trying to push a heavy table: most push in almost the same direction, and one pushes along a direction no one else noticed. The combined motion is dominated by the crowded direction. Orthogonalizing the update changes its geometry so that a few dominant directions do not swallow the others. It balances the singular directions of the update; it does not give every parameter an identical step.

Jordan notes that these update matrices can be nearly low-rank, with a few directions dominating. He speculates that increasing the scale of smaller, "rare" directions is part of why Muon works. That explanation is a hypothesis about weight-update geometry, not proof that Muon detects rare examples in a dataset or guarantees every edge case gets learned. This is where my comparison stops being literal. My proposal concerns which cases receive useful learning signal; Muon concerns how a matrix update is shaped. I care about keeping both questions visible without pretending they are the same mechanism.

Muon is meant for two-dimensional hidden-layer weights; input/output layers and scalar or vector parameters still use AdamW. That is another reason neither optimizer is a universal winner.

Never Give Up: a comparison about who gets attempts

Never Give Up is closer to the other half of my question: a hard case may not produce useful feedback if we give it too few tries. Imagine training a model on two math questions. On the easy one, it gets three of four attempts right. On the hard one, it gets all four wrong. Group-relative RL can compare good and bad attempts on the easy question, so it has something to reinforce. The hard question produces no successful attempt in that small group. It may be exactly where the model needs to learn, yet it offers little usable signal.

Never Give Up: easy problems provide mixed feedback while hard problems may get all wrong attempts.

That is the Matthew Effect the Never Give Up paper studies: already-solvable problems keep producing feedback, while the hardest problems struggle to get a foothold. Giving every problem many more attempts sounds fair, but it also spends extra compute on easy problems that no longer need much help. The paper's answer is adaptive sampling: start with a small group, then give unresolved problems another chance rather than handing all questions the same fixed budget. It uses probabilistic requeueing, not an endless loop until every problem is solved.

In my earlier Never Give Up explainer, I used an example of a problem with only a 2% chance of success per attempt. Four attempts give it about a 7.8% chance of yielding at least one success. Thirty-two attempts raise that chance to about 47.5%. Those are illustrative probabilities, not a reported model accuracy. A hard problem may need more tries before it produces a trajectory RL can learn from. The paper reallocates attempts; it does not implement my proposed reward. I would investigate that choice alongside sampling.

Where I would take the comparison

For Muon, inspect the update, the baselines and the tasks where the optimizer helps or fails. For Never Give Up, inspect which problems get attempts, which ones ever yield a success, and what extra sampling costs. For my approach, inspect the cases behind an aggregate score first: how often they recur, how costly they are when they fail, and whether they are actually learnable. A single top-line number cannot answer any of those questions. The test of my approach is whether tail failures fall without an unacceptable drop on the rest of the task.

I would then test whether rare and hard cases can matter in the reward or training signal without rewarding noise or making the model chase impossible tasks. Pair that with a fairer sampling budget and measure what still fails. Muon suggests that a few loud directions can dominate an update; Never Give Up shows how easy problems can keep producing training feedback while hard ones wait. They are different fixes to different problems. What I take from both is a habit: when the average looks good, ask who has not had a useful chance, and why. The leftover 1% is a map, not automatically a bug or an excuse to stop looking. That is where the next useful idea might be hiding.

Sources: Muon, Keller Jordan · Learning to Solve Hard Problems in RL for LLMs by Never Giving Up · My Never Give Up explainer