Metareasoning

Research on systems that decide what to investigate next, and that learn to make that decision better.

A forecasting system reasons about the world. It also makes a second set of choices: which source to read, which hypothesis to test, which model to fit, and when to stop. Metareasoning is the study of that second set of choices. Its central quantity is the value of a computation, the expected improvement in a later decision that the computation buys, net of what it costs.

The idea is old. Howard priced information in 1966, Matheson priced analysis and computation in 1968, and Russell and Wefald built a theory of metareasoning on the value of computation in 1991. Language-model agents have changed the economics. They propose research actions cheaply and in quantity, so the choice among actions becomes the bottleneck, and their logged histories supply data from which that choice can be learned.

META LEVEL Candidate computations query a source · fit a model test a hypothesis · find a counterexample or stop value of computation expected change in the decision, less cost Run the best one or stop when no value exceeds its cost OBJECT LEVEL Information state what was knowable at time t Forecast and decision a distribution, then an action Outcome resolves later, and scores both resolved outcomes price the computations that preceded them

The meta level chooses which computations run. Resolved outcomes score both levels.

Reasoning answers a question such as “what does this filing imply for next quarter’s revenue?” Metareasoning asks whether the system should read the filing’s notes, pull last year’s comparable, fit an elasticity, look for a disconfirming source, or stop because nothing it could learn would change the decision. The introduction sets out the value of computation and the neighbouring vocabularies: self-improving agents, autonomous research, meta-learning and continual learning.

State of play in September 2026

Research-loop controllers are simple

Agents run the whole cycle of hypothesis, code, experiment and interpretation. AIDE treats machine-learning engineering as a tree search over code, and RD-Agent(Q) chooses its next research direction with a multi-armed bandit scheduler.

Value-of-information reasoning has started to enter these controllers. ExTS treats tree expansion as a value-of-information decision under a small evaluation budget, and AERA learns whether further computation is likely to recover a better answer. Both value computation against a benchmark score. Neither values it against a downstream decision.

For computations run in parallel, the budgeted Brownian race gives the pruning rule. A rollout or an agent is kept while its future pivotal value exceeds its carrying cost. In its committed example the rule beats one-shot screening and static thinning at the same expected budget.

Resolved outcomes as labels

Forecasts resolve, so the passage of time produces supervision that needs no annotator. Future-as-Label trains a model on causally masked information and rewards it with a proper scoring rule once events resolve. It reports a 27% Brier improvement for Qwen3-32B over its pretrained baseline. Training is offline on resolved binary events. The paper states that deployment-time feedback loops “are not evaluated in this study.”

Much of the improvement happens outside the weights

Prompts, memory, tools and control logic can all be updated from experience. GEPA evolves prompts by reflecting on execution traces and reports beating GRPO by 6% on average with up to 35 times fewer rollouts. Against this, Huang et al. find that models struggle to correct their own reasoning without external feedback, and sometimes get worse. Every mechanism with a reported gain relies on an external evaluator.

Reported gains under search-aware evaluation

Systems that generate many candidates and report the best one need search-aware evaluation. FINSABER finds previously reported advantages of LLM investing strategies deteriorate over two decades and more than a hundred symbols. Gençay (2026) deflates every reported result by the agent’s own trial count and rejects every LLM-discovered strategy it tests. The same paper shows a deliberately leaky oracle with a Sharpe ratio of 35 passing both the deflated Sharpe ratio and the probability-of-backtest-overfitting test. Statistical correction does not detect look-ahead.

Look-ahead in language models is measurable

Model families with annual training cutoffs, such as DatedGPT and ChronoGPT, let a backtest use only what a model could have known. A recall-based test estimates the probability that a model has internalised a realised outcome. The effect is not always the expected one: Glasserman and Lin find that in-sample, general knowledge of a company distorts sentiment more than specific knowledge of its later returns.

Further reading