Definition

Test-time compute (TTC) refers to the computational resources – such as tokens, processing time, and memory – that an AI system consumes during inference, as opposed to during model training. Unlike training compute, which is fixed once a model is built, test-time compute is dynamic: it can be scaled up or down in response to task complexity, budget constraints, or quality requirements at the moment a request is processed. The term is closely associated with techniques that deliberately allocate more compute at inference time to improve the accuracy or reliability of AI outputs.

How It Works

Test-time compute scaling operates by dynamically expanding the reasoning or execution resources available to a model when it generates a response. Early implementations include reasoning models that allocate additional tokens to an internal chain-of-thought before producing an answer, and sampling strategies such as best-of-N, where a model generates multiple candidate outputs and the best one is selected. In more advanced agentic systems, TTC orchestration frameworks separate the decision about how much compute to spend – and where – from the LLM’s own reasoning. This allows an orchestration layer to model the expected value of additional compute on a per-step or per-branch basis, allocating resources adaptively across different phases of a complex task rather than applying a uniform compute budget. Execution branches can be run in parallel, validated, and pruned or extended depending on intermediate results.

What It Is Used For

Test-time compute techniques are used to improve the accuracy of AI systems on tasks where a single forward pass through a model is insufficient. They are particularly valuable in agentic workflows involving long-horizon tasks – such as software engineering, legal reasoning, or multi-step data analysis – where errors made early in execution compound over time. TTC scaling allows developers to trade cost for quality in a controlled way, spending more compute on difficult or high-stakes tasks while keeping costs low on straightforward ones. Enterprise AI systems use TTC orchestration to balance accuracy, latency, and cost without requiring retraining or larger models.