TL;DR

AI21 Maestro is a general-purpose agentic framework that automatically scales compute and optimizes orchestration. We demonstrate how it significantly improves model performance on SWE-bench-verified through structured plans, automatic horizontal scaling and decision-theoretic optimization techniques.

Introduction

As agentic tasks grow more complex and the variance in execution increases, the core challenge becomes how to allocate test-time compute (TTC) efficiently across vast and extended trajectories.

Reasoning models were an early attempt to answer this question. These models dynamically allocate more tokens to “think” and better choose their trajectories. Later, strategies like “best of N” (also using more compute to improve accuracy), or “agent/model routing” (learning to route for the cheapest model that solves the problem) joined the agentic game and showed value in many use cases.

While useful, those techniques suffer from two shortcomings: 

  • They rely on “blackbox orchestration”: Current strategies treat each run as a sealed unit rather than part of a coordinated team. These runs act like individuals in separate rooms, agnostic of each other’s progress and only evaluated post-execution.
  • They are trained once, and subsequently applied at inference time according to the strategy with the best average performance on the training set. That means that they not only suffer from the divergence between the distribution of the training set data and that at inference time, but also require full retraining for every new model, agent, prompt, or tool that the system can use. 

Scaling to more powerful models fails to resolve these fundamental flaws, as the underlying limitations are architectural. We argue that addressing these issues requires structured Test-Time Compute mechanisms, specifically, systems that allocate resources explicitly and adaptively during execution to enhance accuracy under resource constraints.

AI21 Maestro is such a general-purpose agentic framework. It separates decision making about orchestration from the LLM reasoning itself, allowing better control over long-horizon, multi-step tasks. To demonstrate its power, and explain some of its workings, we show its impact on SWE-bench tasks. 

Applying AI21 Maestro to SWE-bench

SWE-bench-verified is one of the best known agentic code benchmarks. Its tasks require dozens if not hundreds of steps, which amplify problems such as agents getting stuck or costing a lot to solve a simple problem. Dedicated SWE agents with specific logic and custom tools have dominated the leaderboard for a while, but recent frontier models have caught up and their “naive” execution with a simple bash tool almost matches the tailor-made agents. 

To fairly evaluate AI21 Maestro, we took the following approach: Maestro was only given bash access (and LLMs), and was compared against mini-swe-agent, the ReAct-based agent published by SWE-bench, using the same models and bash tool. We specifically tested AI21 Maestro with GPT-5 and GPT-5 mini, but the same method and gains can be shown for other models. 

Applying AI21 Maestro to SWE-bench
Pink bars are the result of running AI21 Maestro. All the rest are the published results on SWE-bench.

As seen in the above chart, AI21 Maestro significantly improves the performance of GPT-5 and GPT-5-mini. In fact, with Maestro, GPT-5 performs at the general level of Claude 4.5 Opus and Gemini 3 Pro, the top published performers. 

To understand where this improvement comes from, we focus on three components of Maestro (there are others, which we’ll cover elsewhere): 

  • Horizontal scaling
  • Structured plans
  • Exploring the action space

Component 1: Better cost/accuracy with horizontal scaling

As already shown by a few dedicated SWE agents (e.g., TRAE), horizontally scaling cheaper models may be more cost effective than using a stronger, larger model, provided you can reliably identify the best output. In the chart below, this tradeoff is clearly demonstrated: Running 8 trajectories of gpt-5-mini yields a higher score with lower cost as compared to running gpt-5 once. But two challenges remain: (a) scaling automatically and only when it is valuable; and (b) choosing the right trajectory.

Component 1: Better cost/accuracy with horizontal scaling
Percent of resolved items and the median cost per item, across 1, 2, 4 and 8 parallel runs of each model, assuming an oracle reducer that chooses the correct answer. Cost is calculated based on the public OpenAI API pricing.

Maestro solves the first by being built with Test-Time-Compute in mind from the ground up. Expected cost and value is modeled into every “branch” that Maestro considers, based on either learned priors from offline simulation, or manual input from the developer. That way, instead of manually coding parallel execution workflows and optimizing among them, the developer can let Maestro parallelize when possible based on budget.

Whenever Maestro branches out, it keeps full observability and control over those trajectories. Validators are run at the end of the branch, and a Reducer mixes/chooses the final candidate from the branches. This means Maestro can stop agents that get stuck, or just terminate all running branches if one branch finished successfully – without the developer writing any custom code or logic to accomplish this. Maestro’s Execution Engine has built-in support for parallelizing even with state-changing actions, whereas naive parallel execution would otherwise cause conflicting writes and inconsistent state. This approach alone brings immense benefits and potential, as the scores show.

Component 2: The power of structured plans 

Most general-purpose agents today are descendants of the ReAct framework: agents built around raw prompting and ReAct-style loops. Planning, decision-making, and execution are mediated through token generation, with choices such as when to continue, retry, branch, or stop being inferred from next-token predictions. The model encodes its entire control flow inside its textual context, which makes the agent look like a program, but behave like a stochastic process with high variance in cost, latency, and accuracy.

Component 2: The power of structured plans
The cost of each item in the dataset was measured across 16 runs, and coefficient of variation (CV = std/mean) was calculated for each item. The boxplot shows the distribution of CV across the dataset. The variance is high with similar distributions for GPT-5 and GPT-5 mini both for successful and failed runs.

Let’s look at the prompt of the ReACT agent defined in mini-swe-agent. It’s a simple prompt, outlining the basic steps of solving any SWE task: 

## Recommended Workflow

This workflows should be done step-by-step so that you can iterate on your changes and any possible problems.

1. Analyze the codebase by finding and reading relevant files
2. Create a script to reproduce the issue
3. Edit the source code to resolve the issue
4. Verify your fix works by running your script again
5. Test edge cases to ensure your fix is robust
6. Submit your changes and finish your work by issuing the following command: `echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT`. Do not combine it with any other command. <important>After this command, you cannot continue working on this task.</important>

Looking at this from a TTC angle, not all steps are born equal. Some are probably more complicated than others and would benefit from using more resources. Some are read-only, and are therefore easier to parallelize, while others change the state of the repository. In contrast with blackbox ReAct-type agents, AI21 Maestro uses structured programmatic plans written in AI21 Maestro’s Plan Definition Language, which enable parallel execution of independent steps and step-level dynamic TTCS (Test-Time Compute Scaling).

Blackbox TTCS vs Structured

In order to smartly orchestrate those trajectories, Maestro uses structured programmatic plans written in AI21 Maestro’s Plan Definition Language. These plans, written either by the user or dynamically by a Planner model, represent a logical decomposition of the task into discrete steps while abstracting away the concrete orchestration and TTCS decisions to be made within each step. They replace text-based reasoning/ReACT trajectories with code that can be compiled, parsed and orchestrated. 

In addition to enabling parallelism and step-level TTCS, Maestro’s Structured Plans offer other key benefits over ReAct / Reasoning models agentic plans:

  1. Explicit control and data flow: Written in code, Maestro’s structured plans can dictate data dependencies, control flow, variables, conditions and other “glue code” – things free text simply can’t achieve.
  2. Better explainability: Instead of trying to debug a 40K-long reasoning chain, plans are short and encoded in hierarchical code that can be easily understood.
  3. Plan enforcement: Reasoning models and ReAct agents tend to get stuck or diverge from their original plan. With structured plans, execution can be enforced and controlled even in long trajectories. 
  4. Python native: Unlike with some other frameworks with proprietary flow representations, since modern LLMs are quite proficient in Python, models are able to produce great PDL plans even without dedicated training. 
  5. Secured by design: Since models can create plans at runtime, PDL is limited to the minimum syntax that’s required to define plans. Model-generated plans can’t import libraries and perform other insecure operations. 

For SWE-bench, we evaluated two variants of structured plans: The first, a user-created plan that mimics the same workflow detailed in the mini-swe-agent prompt; The second, a planner model that dynamically plans for each item at runtime.

Examples of two Structured plans: on the left a user-defined swe-agent prompt, on the right a part of a dynamically generated plan.
Examples of the two Structured plans methods used: on the left, a user-defined structured version of the same plan used in mini-swe-agent’s prompt for the entire benchmark. On the right, 2 examples of a part of a dynamically generated plan by a Planner model.

Both methods show superiority compared to the ReAct-based agent with the same model, even in a single trajectory. As we scale those horizontally, we see that the structured plans improve at a higher rate.

The difference at runtime can be seen when sampling an item from the dataset and looking at the trajectories. In the chart below, we see that the ReAct agent reaches an early false “aha” moment and commits to patching before building sufficient awareness of the issue. It then iterates between fixing and testing without reaching a resolution. In contrast, the structured execution spent more effort on building context and writing tests ahead of time, which eventually made it succeed.
Percent of resolved items across 1, 2, 4 and 8 parallel runs of each planning method, on top of GPT-5 mini, assuming an oracle reducer that chooses the correct answer.

The difference at runtime can be seen when sampling an item from the dataset and looking at the trajectories. In the chart below, we see that the ReAct agent reaches an early false “aha” moment and commits to patching before building sufficient awareness of the issue. It then iterates between fixing and testing without reaching a resolution. In contrast, the structured execution spent more effort on building context and writing tests ahead of time, which eventually made it succeed. 

The difference at runtime can be seen when sampling an item from the dataset and looking at the trajectories. In the chart below, we see that the ReAct agent reaches an early false “aha” moment and commits to patching before building sufficient awareness of the issue. It then iterates between fixing and testing without reaching a resolution. In contrast, the structured execution spent more effort on building context and writing tests ahead of time, which eventually made it succeed.
On the left, a breakdown of the different steps across the major plan parts, on an example sample from the benchmark. The ReAct steps were manually allocated by reading its outputs. On the right side, the overall distribution of allocated steps across the major plan parts for the predefined plan across the entire benchmark.

Component 3: Exploring the action space

Building on the pillars discussed so far, Maestro now has a range of actions it can perform at any given moment. Choosing agents that span different models with different cost profiles (often run in parallel), planning and plan-based trajectories using multiple models, and mechanisms to validate, prune, or halt execution paths as they unfold. The challenge is how to optimally navigate this action space given the goal and resources, deciding which actions to take, and when.

Rather than optimize for a single “best” average action, AI21 Maestro assembles a portfolio of complementary actions that collectively span the problem space. In practice, a diverse set of medium-performing agents, each strong in different areas, outperforms a small number of state-of-the-art agents that converge on the same solutions. The real challenge is selecting and orchestrating among them efficiently.

Done properly, this should close the gap between what a single agent can achieve and what an optimal orchestrator would choose: The cheapest execution that succeeds for each task. Instead of scaling models or thinking tokens, this approach shifts the Pareto frontier by using existing capabilities more effectively.

Each grey line represents the accuracy of a single trajectory (including different GPT-5 family models, Maestro structured, constrained by a maximum cost in dollars per item in the dataset (x-axis). The pink line represents the accuracy of an optimal orchestrator choosing the best possible action for any given cost.

Each grey line represents the accuracy of a single trajectory (including different GPT-5 family models, Maestro structured, constrained by a maximum cost in dollars per item in the dataset (x-axis). The pink line represents the accuracy of an optimal orchestrator choosing the best possible action for any given cost. 

Summary

AI21 Maestro is an orchestration framework that offers automated, fine-grained optimization of long-horizon agentic flows. It is general-purpose and applies broadly, but to illustrate its power we showed how it boosts the performance on SWE-bench. We also described some of the elements of Maestro that enable this performance, including horizontal scaling and automated smart parallelism, the use of Python-based structured plans, and exploring the huge action space using decision-theoretic optimization techniques.