← All writing
AI AlignmentPaper deep dive

Can one reward teach an assistant everything?

A reading of MAHALO, where accuracy, human values and tutoring pull an assistant in different directions.

Dhruvi Paprunia7 min read
READING NOTES / DHRUVI PAPRUNIACan one reward teachan assistant everything?One tutoring question. Three goals that do not always agree.“Help me solve this, but let me try.”CORRECTIs the math right?HELPFULDoes the step teach?ENGAGINGWill they keep trying?THREE QUESTIONS. NO SINGLE EASY SCORE.Conceptual map, not a measured result from the paper.

Imagine asking an AI tutor for help with a math problem. It could give the right answer immediately, but that may not help the student learn. It could ask a thoughtful question instead, but the question is no use if it leads the student in the wrong direction. Correctness and good teaching overlap, yet they are not the same thing.

Training a model often involves a reward: a score that tells it which responses to favor. A single score can hide the difference between goals. What if an assistant needs to be accurate, honest, and engaging, and those aims do not always agree? That is the question Yiran Shen, Yu Xia, Jonathan Chang, and Prithviraj Ammanabrolu take up in their paper on training for several goals at once.

Three kinds of feedback

One prompt, different questions
MATHCan we check it?
VALUESWhich answer helps?
TUTORINGWhat happens next?
Conceptual reading of the paper, not a measured result.

Some outcomes have an answer we can check. In math, a result can be compared with a known solution. The paper calls this a verifiable setting. Other judgments depend on people. Two answers may both be correct, but readers may prefer the more honest or helpful one. A tutoring exchange adds another complication: what happens across several turns matters more than one reply on its own.

Think again about the tutor. "The answer is 12" may pass a correctness check. "What happens if you divide both sides by three?" may help a student take the next step. A cheerful response that quietly teaches a false rule may score well for engagement but badly for accuracy. These are examples of why the goals can conflict, not evidence that this particular system has solved tutoring.

The authors call their approach MAHALO. It keeps several goals visible rather than folding them into one permanent score. Most of the model is shared: this common part, or backbone, handles language for all the goals. Smaller output parts called action heads learn different preferences. Picture one assistant with several ways to shape its next response, rather than three entirely separate assistants.

Keep goals visible as the response unfolds
SHARED MODELLanguage
GOAL HEADSDifferent aims
STEP SCORENext choice
Conceptual reading of the paper, not a measured result.

During training, the model sees pairs of possible responses and learns which response is preferred for each goal. This comparison method is called Direct Preference Optimization, or DPO. For a tutoring prompt, one pair might show a response that nudges the student and another that simply gives away the answer. The preference depends on what is being taught, so the system keeps that goal explicit.

The framework also uses process reward models. These are models trained to score a step along the way, rather than judging only the final answer. If an assistant is solving a problem in several steps, such a model can assess a candidate next step before the assistant continues. For example, a correct final number does not tell us whether the first algebra move was sound. A step score can offer a different signal, though it can still make mistakes.

At response time, the system can give the goals different weights and compare possible next steps. Response time is often called inference: the model is being used, not trained. A user who wants a careful hint rather than an immediate answer needs a different balance from someone checking a result. Changing a weight is a way to steer that balance, not a guarantee that every response will match the user's intent.

What can be checked, and what has to be judged?

The paper tests math reasoning, human values, and multi-turn tutoring. Its results do not say that the same method wins in every setting. A checker can give a fairly clear signal for a math answer. Whether an explanation is helpful or a tutoring turn is engaging needs a judgment about the response and its context.

The chart below is from the paper's math setting. It changes the balance between two trained goals, accuracy and engagement, and plots both scores. More weight on one goal does not turn into a clean improvement in both. The bars around points show variation, so it would be too much to claim a precise winning setting from this chart alone. Its useful lesson is that the choice of balance is visible and testable.

Figure 2 from the MAHALO paper showing math accuracy and engagement across goal weights
Figure 2 from Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards. Math accuracy and engagement under different goal weights. The vertical scales differ.

The authors report that searching among candidate steps with a process score helped most in the checkable math setting they tested. For less checkable preferences, training with several goals mattered more. This is a result from their experiments, not a rule for all assistants.

That leaves a question beneath the scores: who decided what a good step looks like? A step scorer trained on incomplete criteria might favor a plausible explanation that is quietly wrong. The reward is not a neutral definition of quality. It is a signal built from examples and judgments, and it needs its own evaluation.

What I would want to see next

The tutoring case is where the separate goals become hard to fake. Give the assistant a student who has made a specific algebra mistake. Compare a reply that fixes the answer for them with one that asks a useful question. Can the student explain the corrected step and solve a similar problem without another hint? That tests engagement and accuracy against what the interaction actually teaches, not just whether a reply sounds helpful.

Then change what the student needs halfway through the conversation. They begin by wanting a hint, but later ask for the answer because time is running out. Can the assistant shift its goal weights without losing track of the earlier mistake? This connects MAHALO to Ammanabrolu's work on agents that learn from human feedback while acting with people. The key check is what the student does next. If a judge likes the assistant's wording but the student repeats the mistake, the step score has missed the thing tutoring was meant to improve. That is a test I would run, not a result this paper reports.

The point is not to find one perfect setting for a "good assistant." It is to see what is gained and lost when its goals compete. Keeping those goals separate is what lets us ask that question in the first place.

Paper: Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards