← All writing
AI AgentsPaper deep dive

The agent did everything right. The website still said no.

What a verifier can learn when the assistant took the right steps but the website failed.

Dhruvi Paprunia8 min read
READING NOTES / DHRUVI PAPRUNIAThe agent did everything right.The website still said no.A check that looks beyond the last screenA TRAIN TICKET, ALMOST BOOKED✓ Train selected✓ Passenger details filled✓ Travel day checked! Verification screen failed · Ticket not bookedPROCESS ≠ OUTCOME

Suppose an assistant is booking a train ticket. It finds the right train, fills the passenger details, then hits a verification screen that never loads. The ticket is not booked. Did the assistant fail? Now imagine it books a ticket but chooses the wrong day. A simple "done" check could reward the second attempt and punish the first. That is a problem if the check is used to decide which assistant to trust or what it should learn next.

Corby Rosset, Pratyusha Sharma, Andrew Zhao, Miguel Gonzalez-Fernandez, and Ahmed Awadallah study this problem in The Art of Building Verifiers for Computer Use Agents. A verifier is a system that inspects an assistant's actions and decides whether it completed the task. Their Universal Verifier tries to make that decision from evidence, not from the assistant saying it succeeded.

Two questions instead of one

A browser task has a goal and a trail of actions. The paper calls that trail a trajectory: the screens the assistant saw, what it clicked or typed, and what changed afterward. The verifier separates the outcome, whether the goal happened, from the process, whether the assistant made sensible choices along the way.

Both attempts can end with the same red "not booked" result, yet only one is the assistant's mistake. In the first, the assistant enters the correct details and the website's verification screen fails. In the second, it chooses the wrong travel day. That difference changes what we should teach the assistant. The verifier asks whether an error was within its control; if the website failed first, it should not count every later blocked step as another independent mistake.

Same result. Different cause.
A / Website failsCorrect train → Correct day → Verification screen failsNOT BOOKED · website caused it
B / Agent errsCorrect train → Wrong day → SubmissionNOT BOOKED · agent caused it

To tell those two stories apart, the verifier needs more than a verdict. A second problem is what it gets to see. If it only reads the last screen, it may miss the brief "payment declined" message that appeared earlier, or a date that changed just before submission. The authors select relevant screenshots for each task criterion so the judge can inspect evidence from across the attempt. For a train booking, one criterion might be the travel day and another the destination. The criteria should be distinct enough that one mistake is not counted twice.

What the paper actually tested

That is the proposed method. Did it work better than a simple judge? The team built agent attempts and asked people to judge whether each goal actually happened. The verifier and two comparison judges then tried to match those human answers. But a high match rate by itself can be a trick.

Imagine 100 attempts where 90 really failed and 10 succeeded. A lazy judge says "failed" for every one. It agrees with the humans on all 90 failures, so its raw agreement is 90%. Yet it misses every successful booking. Because 90 of the 100 outcomes are failures, this judge can get 90 agreements by repeating the common answer. After correcting for that lucky baseline, its Cohen's kappa score is zero. This is an invented teaching example, not the paper's result.

The "always failed" trap

90 failed + 10 succeeded. A judge says "failed" 100 times.

90% raw agreement. Zero skill at finding success.
Invented teaching example: kappa removes agreement earned by repeating the common label.

That is why the paper reports Cohen's kappa rather than raw agreement alone. It asks: how much did the judge agree with humans beyond what their yes/no habits would produce by chance? Take a less lopsided example: out of 100 judgments, a human and a verifier agree on 80, so observed agreement is 0.80. Their label habits would yield 50 agreements by chance, or 0.50. Subtract that chance agreement, then divide by the agreement still available: (0.80 - 0.50) / (1 - 0.50) = 0.60. Again, these 100 judgments illustrate the calculation; they are not paper data.

How kappa discounts lucky agreement

100 judged attempts · 80 agreements · 50 expected by chance

(0.80 - 0.50) / (1 - 0.50) = 0.60
Illustrative calculation, not a 100-attempt result from the paper.

With the math in mind, we can read the actual comparisons. On the paper's first labeled set of 140 trajectories, the Universal Verifier reached a kappa of 0.64 against human outcome judgments. The two baseline judges scored 0.31 and 0.44. On a second set of 106 trajectories, its score was 0.58 against 0.13 and 0.26. The compact table compares agreement beyond chance, not task success.

Agreement with human outcome labels
Test setVerifierBaseline 1Baseline 2
140 attempts0.640.310.44
106 attempts0.580.130.26
Cohen’s kappa, not percent correct or task completion.

A kappa of 0.64 does not mean the verifier was correct on 64% of attempts. Nor does it mean 64% of tickets were booked. Those would be different questions.

Agreement is one test. The dangerous direction of disagreement is a false positive: calling an unsuccessful task successful, the kind of mistake that could send someone a false "booked" notice. The paper reports false-positive rates of 0.01 and 0.08 on its two labeled sets. To picture the rate, imagine 100 genuinely failed attempts in each setting: the verifier would wrongly say "done" for about 1 in the first and 8 in the second. That is a rate illustration, not the count of mistakes in the paper's 140 and 106 trajectories. These rates are low in these tests, not a guarantee for a new website. Giving baseline judges a stronger model did not erase the gap, which supports the authors' point that evidence and scoring design matter alongside the model.

Imagine 100 tasks that actually failed

First test: 1 falsely called done

Second test: 8 falsely called done

Illustration of 0.01 and 0.08 false-positive rates, not the paper’s raw counts.

The next question was not only how the verifier scored, but how its design was found. The paper also asked whether an automated research agent could build one from scratch. Its many quick edits improved the score, but missed several large structural changes the human designer made. The original paper's chart below tracks agreement as those experiments progressed. It is not a claim that people always design better verifiers. Here, the question was whether the automated search found the same useful structure under this setup.

Paper Figure 1, agreement with human labels during verifier design experiments
Figure 1 from The Art of Building Verifiers for Computer Use Agents. Agreement with human outcome labels during human expert and automated verifier-design experiments. The marked blue jumps were structural human decisions.

What I would test next

That leaves a question the averages cannot answer: can it name why a particular attempt failed? The next test I want is a set of paired agent traces that end on the same screen for different reasons. In one, the assistant makes a wrong choice early. In the other, the website blocks a correct path. Show a verifier the final screen, then the full trail, and ask it to name the decisive step and whether it was controllable. If its verdict changes for the right reason, we learn more than we do from another average agreement score.

The paper treats giving up after one failed attempt as insufficient effort, but it does not say how much retrying is enough. I would test the same task against a page that fails once and then works, and one that stays broken. Give the agent a limited retry budget and ask the verifier to score its recovery choices: when it retries, when it switches approach, and when it stops with an honest account of what happened. A useful process score should not reward endless clicking on a dead page, or punish an agent for declining to pretend that a failed booking succeeded.

I would then use those verdicts as training feedback for an agent. Does a process score help it recover from a real mistake without teaching it to avoid tasks whenever a site is flaky? Sharma's work on reasoning and sequential decisions makes that the important next link: the label is only useful if it teaches better behavior on the next attempt. I would test on changed website layouts and have people review every case the verifier calls "done" but the site state does not support. This is a proposed experiment, not one the paper reports.

A better judge cannot make a broken website work. It can tell us whether the assistant caused the failure, and keep a false "done" from looking like progress. Cheap, confident judges with escalation, like the JEV cascade I wrote about in my series, are one way to get better evidence without a bigger judge.

Paper: The Art of Building Verifiers for Computer Use Agents