The agent did everything right. The website still said no.
What a verifier can learn when the assistant took the right steps but the website failed.
Suppose an assistant is booking a train ticket. It finds the right train, fills the passenger details, then hits a verification screen that never loads. The ticket is not booked. Did the assistant fail? Now imagine it books a ticket but chooses the wrong day. A simple "done" check could reward the second attempt and punish the first. That is a problem if the check is used to decide which assistant to trust or what it should learn next.
Corby Rosset, Pratyusha Sharma, Andrew Zhao, Miguel Gonzalez-Fernandez, and Ahmed Awadallah study this problem in The Art of Building Verifiers for Computer Use Agents. A verifier is a system that inspects an assistant's actions and decides whether it completed the task. Their Universal Verifier tries to make that decision from evidence, not from the assistant saying it succeeded.
Two questions instead of one
A browser task has a goal and a trail of actions. The paper calls that trail a trajectory: the screens the assistant saw, what it clicked or typed, and what changed afterward. The verifier separates the outcome, whether the goal happened, from the process, whether the assistant made sensible choices along the way.
Both attempts can end with the same red "not booked" result, yet only one is the assistant's mistake. In the first, the assistant enters the correct details and the website's verification screen fails. In the second, it chooses the wrong travel day. That difference changes what we should teach the assistant. The verifier asks whether an error was within its control; if the website failed first, it should not count every later blocked step as another independent mistake.
To tell those two stories apart, the verifier needs more than a verdict. A second problem is what it gets to see. If it only reads the last screen, it may miss the brief "payment declined" message that appeared earlier, or a date that changed just before submission. The authors select relevant screenshots for each task criterion so the judge can inspect evidence from across the attempt. For a train booking, one criterion might be the travel day and another the destination. The criteria should be distinct enough that one mistake is not counted twice.
What the paper actually tested
That is the proposed method. Did it work better than a simple judge? The team built agent attempts and asked people to judge whether each goal actually happened. The verifier and two comparison judges then tried to match those human answers. But a high match rate by itself can be a trick.
Imagine 100 attempts where 90 really failed and 10 succeeded. A lazy judge says "failed" for every one. It agrees with the humans on all 90 failures, so its raw agreement is 90%. Yet it misses every successful booking. Because 90 of the 100 outcomes are failures, this judge can get 90 agreements by repeating the common answer. After correcting for that lucky baseline, its Cohen's kappa score is zero. This is an invented teaching example, not the paper's result.
90 failed + 10 succeeded. A judge says "failed" 100 times.
90% raw agreement. Zero skill at finding success.That is why the paper reports Cohen's kappa rather than raw agreement alone. It asks: how much did the judge agree with humans beyond what their yes/no habits would produce by chance? Take a less lopsided example: out of 100 judgments, a human and a verifier agree on 80, so observed agreement is 0.80. Their label habits would yield 50 agreements by chance, or 0.50. Subtract that chance agreement, then divide by the agreement still available: (0.80 - 0.50) / (1 - 0.50) = 0.60. Again, these 100 judgments illustrate the calculation; they are not paper data.
100 judged attempts · 80 agreements · 50 expected by chance
(0.80 - 0.50) / (1 - 0.50) = 0.60With the math in mind, we can read the actual comparisons. On the paper's first labeled set of 140 trajectories, the Universal Verifier reached a kappa of 0.64 against human outcome judgments. The two baseline judges scored 0.31 and 0.44. On a second set of 106 trajectories, its score was 0.58 against 0.13 and 0.26. The compact table compares agreement beyond chance, not task success.
A kappa of 0.64 does not mean the verifier was correct on 64% of attempts. Nor does it mean 64% of tickets were booked. Those would be different questions.
Agreement is one test. The dangerous direction of disagreement is a false positive: calling an unsuccessful task successful, the kind of mistake that could send someone a false "booked" notice. The paper reports false-positive rates of 0.01 and 0.08 on its two labeled sets. To picture the rate, imagine 100 genuinely failed attempts in each setting: the verifier would wrongly say "done" for about 1 in the first and 8 in the second. That is a rate illustration, not the count of mistakes in the paper's 140 and 106 trajectories. These rates are low in these tests, not a guarantee for a new website. Giving baseline judges a stronger model did not erase the gap, which supports the authors' point that evidence and scoring design matter alongside the model.
First test: 1 falsely called done
Second test: 8 falsely called done
The next question was not only how the verifier scored, but how its design was found. The paper also asked whether an automated research agent could build one from scratch. Its many quick edits improved the score, but missed several large structural changes the human designer made. The original paper's chart below tracks agreement as those experiments progressed. It is not a claim that people always design better verifiers. Here, the question was whether the automated search found the same useful structure under this setup.
What I would test next
That leaves a question the averages cannot answer: can it name why a particular attempt failed? The next test I want is a set of paired agent traces that end on the same screen for different reasons. In one, the assistant makes a wrong choice early. In the other, the website blocks a correct path. Show a verifier the final screen, then the full trail, and ask it to name the decisive step and whether it was controllable. If its verdict changes for the right reason, we learn more than we do from another average agreement score.
The paper treats giving up after one failed attempt as insufficient effort, but it does not say how much retrying is enough. I would test the same task against a page that fails once and then works, and one that stays broken. Give the agent a limited retry budget and ask the verifier to score its recovery choices: when it retries, when it switches approach, and when it stops with an honest account of what happened. A useful process score should not reward endless clicking on a dead page, or punish an agent for declining to pretend that a failed booking succeeded.
I would then use those verdicts as training feedback for an agent. Does a process score help it recover from a real mistake without teaching it to avoid tasks whenever a site is flaky? Sharma's work on reasoning and sequential decisions makes that the important next link: the label is only useful if it teaches better behavior on the next attempt. I would test on changed website layouts and have people review every case the verifier calls "done" but the site state does not support. This is a proposed experiment, not one the paper reports.
A better judge cannot make a broken website work. It can tell us whether the assistant caused the failure, and keep a false "done" from looking like progress. Cheap, confident judges with escalation, like the JEV cascade I wrote about in my series, are one way to get better evidence without a bigger judge.
Paper: The Art of Building Verifiers for Computer Use Agents