In Brief

Before spending more compute on your agent, ask this first: Is your agent able to verify its own outcome (i.e., can your agent recognize when it has the right answer)? For some tasks, verifiability is straightforward; the task hands you conditions to check candidates against and make a stopping decision. For others, it’s less clear. In this post, we discuss why that single property – verifiability – should dictate where you spend your budget next: on verification when completion is checkable, and on diversity and aggregation when it isn’t.

The case against scaling first

Poll most engineers about how to improve an agent and most will say: scale it. More attempts, longer and enhanced context, a stronger model. It assumes the current architecture is already right – just underpowered. Most of the time, scaling works. But it’s the expensive way to go.

Over a year of agent optimization research, we have seen over and over again why this assumption (right architecture, not enough power) is wrong. The giveaway is discovering inefficiencies in the architecture: applying a uniform budget to variable tasks, generating signals and rollouts only to discard them, running a frontier model when an open model would do just as well. Each of these moments indicates there’s room to further optimize the agent’s architecture before scaling.

There’s a simple litmus test: can a candidate be checked against anything? That’s a property of the task, and it forks your architecture. If yes, invest in verification and selection. If no, invest in diversity and aggregation.

Then a second question, about your pipeline rather than the task: is it already exploiting the check available to it? That one we answer with an oracle experiment.

Figure 1. Decision framework for agent architecture. The governing question is whether a candidate answer can be verified, a property of the task rather than of the model. Verifiable tasks favor investment in verification and selection; unverifiable tasks favor diversity and aggregation.

A verification mechanism buys you one thing: permission to commit to a single candidate and throw the rest away. When you can’t build one, diversity and aggregation are the proxy: run varied attempts, then combine them, because without a check you can’t safely commit to any single one.

In this post, we use this litmus test to examine agent pipelines across four domains – agentic search, deep research, RAG indexing, and agentic coding – using our answers to optimize existing architecture and lift performance on each one.

Can the agent tell when it has the right answer? Yes. The answer is a single entity, and the conditions are explicit.

A multi-hop agentic search question specifies its answer through conditions, and the answer is a single short entity, such as one person. That shape is what makes checking efficient. Verifying one entity against stated conditions covers a far narrower space than the original search did.

Change the shape and the answer changes with it. Ask instead for every Nobel laureate whose mother didn’t finish high school, and each candidate you find is still checkable, but knowing you’ve found them all isn’t, because confirming completeness means re-running the search. Precision stays verifiable, coverage doesn’t. That same split returns in Task #4.

Before building anything, we ran an oracle experiment: we made the selection based on the ground-truth answers, creating a ceiling no runtime system could beat, since no real system actually gets to see them. It’s easy to do – and it disqualifies strategies fast. For instance, we saw that pass@k sits far above pass@1, suggesting that a correct answer routinely shows up across many attempts, yet any one single attempt usually won’t be the one to find it. 

Knowing that the answers were already in the pool – just not chosen – helped us direct our focus to verification. We demonstrate two ways to improve verification signals, paying only when quality warrants it:

The free signal waiting to be utilized. On BrowseComp-Plus, models’ own confidence scores track accuracy well enough, once calibrated, to make a selection. Paired with an ensemble of models that fail on different queries, this reached 95.18% accuracy and state of the art.

The signal worth paying for. Confidence is self-assessment, so it adds no new evidence about the world, and models are perfectly capable of being wrong and confident at the same time. That’s tolerable when generators are strong and well calibrated. 

We saw that when testing ensembles of small models on FACTS-Search. Turns out, small models frequently produce wrong answers – a problem majority voting can’t overturn as long as the wrong answer is a popular one. So instead, we swapped confidence scores with an independent verifier that re-researches each candidate from scratch and vetoes the wrong ones.

Training our own 8B verifier, we matched the quality of a frontier verifier on an open + closed model ensemble – at a 3.2x lower cost. Applying the same verifier to an all-open-source ensemble, it lifted accuracy 28% over plain majority voting on identical candidates – more than 80% of what a frontier verifier got on that pool, for roughly 1/100th the price.

Figure 2. Majority voting versus verification prior to voting, on a single query. Six sampled attempts yield three distinct answers (A ×3, B ×2, C ×1), of which only C is correct. Majority voting selects A. An independent verifier rejects A and B before aggregation, and voting over the surviving candidates selects C.

Task #2: Deep Research

Can the agent tell when it has the right answer? No. There is no way to know what a complete report should contain.

A deep research task asks an agent to investigate an open question and produce a report. What makes a report good is coverage: the set of relevant facts, findings and caveats it surfaces. Nobody can write that set down in advance. Compare Task #1, where the question’s own conditions were the enumeration, here the target is a set no one can enumerate before the work is done.

That splits the task in two. Any individual claim in the report is checkable, the same way each name on that Nobel list was checkable. Whether the report is complete is not, because confirming completeness means doing the research over again. Precision survives, coverage doesn’t, and coverage is what a report is judged on.

DeepResearch Bench II makes this concrete: It grades reports against a long list of expert-written rubric items, heavily weighted toward specific facts the answer should contain. The items stay hidden until grading, but the hiding isn’t the root problem. In production you’d hand the model the rubrics gladly. The point is that the list had to be built by experts after the fact, because there was no way to produce it up front. In production there’s no list at all.

So no verifier can rank one report above another on a metric that matters. Diversity and aggregation would have to be our strategy here, and that’s the point: they aren’t an alternative to verification, they’re a fallback.

So we ran some oracle experiments, this time with two strategies: 

  • Selection, the move that worked in Task #1, means keeping the single best report and discarding the rest. Even with an oracle selector, the ceiling sat below the state of the art, which is the earlier finding arriving as a number: without a check, committing to one candidate loses.
  • Merging means taking several reports and combining their facts into one; with an oracle merger, the ceiling cleared SOTA comfortably.
Figure 3. Rubric coverage across reports on DeepResearch Bench II (illustrative). Columns denote reports from individual mid-ranked agents; rows denote rubric items; filled cells indicate coverage. Each report covers a minority of items, and coverage is weakly correlated across reports, so the merged report covers nearly all items while no single report approaches it. This is consistent with the oracle result: selection is bounded below the state of the art, merging is not.

That gap between the two scores – from merging and published SOTA – indicated the right facts were already there – just scattered.

So we built a pipeline that would assemble them all in one place. Using an algorithm we built, we merged reports across seven low-ranked agents from the leaderboard, so one report inherited the facts from all of them. That merged report took first place, above every agent that produced it and above the previous state of the art.

Task #3: RAG Indexing

Can the agent tell when it has the right answer? No, and it doesn’t yet know what will be asked.

Indexing presents perhaps one of the largest challenges to verifiability because the decision – what chunk size? – has to be made in advance, before any question has been posed. A policy handbook has to serve both “what’s the exact notice period in section 4.2” and “can I expense dinner with a client.” The first is answered by a single line. The second is answered only by several rules at the other end of the document, none of which is the answer on its own. Committing to a single chunk size leaves 20-40% of recall on the table.

Figure 4. Query-dependent chunk granularity. Two queries over the same document require retrieval units of different size: a single line for the first, several subsections for the second. Since chunk size is fixed at indexing time, before any query arrives, no single granularity serves both.

This is an instance where we need diversity. Yes, indexing at multiple resolutions costs more storage and more fan-out, but it’s the right trade (can you imagine a support assistant forgoing retrieving a right answer to save on cost?). Inability to verify doesn’t always mean spending less.

Task #4: Agentic Coding

Can the agent tell when it has the right answer? At the end yes, in the middle no.

Verification for a coding task is complex, which is why it’s both the most useful case and the most commonly misinterpreted.

At the end of a run, the check is as hard as checks get: tests pass or they don’t, and you can’t merge two patches into a better patch. During the run, there’s nothing. You could reach most of a repository and still not be able to rule out files. 

So even though the end conditions are binary, the interior is coverage-bound – making coding agents look a whole lot like recall problems (or a deep research report). We saw this in our experiments: feeding candidate patches into context search, rather than only the issue description, made the pipeline far better at finding every relevant file. When we did that, the share of issues where the agent missed all relevant files fell from 9.3% to 2.7%, buying 3.2 points in resolve rate. A large recall gain earning only a modest accuracy gain is the fingerprint of an unverifiable step being treated as verifiable.

To match each step to its verifiability potential, we staffed them each differently. Junior models (open-source MiniMax-M3) fan out across the repo in parallel, covering what nobody can confirm is covered. A senior model (GPT-5.2) distills their attempts into what the fix depends on. Only then does the principal, a frontier model, write the patch once, at the single moment a genuine check exists: 80.8% resolve rate on SWE-Bench Pro, state of the art, at $5.99 per task against $18.28 for a solo frontier agent that scores lower.

Figure 5. Coding pipeline partitioned by verifiability. Junior open-source models explore the repository in parallel; a senior model distills their trajectories into a brief of relevant functions, callers and tests; a principal frontier model generates the patch once. Frontier inference is confined to the only stage with a correctness signal, the test suite.

Summary

Every one of these started with a cheap measurement rather than a bigger model: read pass@k against pass@1, price a verifier against a generator, count the rubrics, run an oracle experiment. Models will keep improving, and that won’t change this: whether an agent can recognize its own success is a property of the task, not of the model running it. Ask that first. Then decide whether you still need to scale.

Acknowledgements

This blog draws on research by Adi Elbaz, Eran Goldstein, Tamar Levy Loboda, Niv Granot, Assaf Gerner, Eli Lepkifker, Oded Avraham, Guy Rephaeli, Guy Freund, Alan Arazi, Ido Weiss and Roee Hendel. Huge thanks to Joanna Kramer for editing.