Metareasoning
Research on systems that decide what to investigate next, and that learn to make that decision better.
A forecasting system reasons about the world. It also makes a second set of choices: which source to read, which hypothesis to test, which model to fit, and when to stop. Metareasoning is the study of that second set of choices. Its central quantity is the value of a computation, the expected improvement in a later decision that the computation buys, net of what it costs.
The idea is old. Howard priced information in 1966, Matheson priced analysis and computation in 1968, and Russell and Wefald built a theory of metareasoning on the value of computation in 1991. Language-model agents have changed the economics. They propose research actions cheaply and in quantity, so the choice among actions becomes the bottleneck, and their logged histories supply data from which that choice can be learned.
The meta level chooses which computations run. Resolved outcomes score both levels.
Reasoning answers a question such as “what does this filing imply for next quarter’s revenue?” Metareasoning asks whether the system should read the filing’s notes, pull last year’s comparable, fit an elasticity, look for a disconfirming source, or stop because nothing it could learn would change the decision. The introduction sets out the value of computation and the neighbouring vocabularies: self-improving agents, autonomous research, meta-learning and continual learning.
State of play in September 2026
Research-loop controllers are simple
Agents run the whole cycle of hypothesis, code, experiment and interpretation. AIDE treats machine-learning engineering as a tree search over code, and RD-Agent(Q) chooses its next research direction with a multi-armed bandit scheduler.
Value-of-information reasoning has started to enter these controllers. ExTS treats tree expansion as a value-of-information decision under a small evaluation budget, and AERA learns whether further computation is likely to recover a better answer. Both value computation against a benchmark score. Neither values it against a downstream decision.
For computations run in parallel, the budgeted Brownian race gives the pruning rule. A rollout or an agent is kept while its future pivotal value exceeds its carrying cost. In its committed example the rule beats one-shot screening and static thinning at the same expected budget.
Resolved outcomes as labels
Forecasts resolve, so the passage of time produces supervision that needs no annotator. Future-as-Label trains a model on causally masked information and rewards it with a proper scoring rule once events resolve. It reports a 27% Brier improvement for Qwen3-32B over its pretrained baseline. Training is offline on resolved binary events. The paper states that deployment-time feedback loops “are not evaluated in this study.”
Much of the improvement happens outside the weights
Prompts, memory, tools and control logic can all be updated from experience. GEPA evolves prompts by reflecting on execution traces and reports beating GRPO by 6% on average with up to 35 times fewer rollouts. Against this, Huang et al. find that models struggle to correct their own reasoning without external feedback, and sometimes get worse. Every mechanism with a reported gain relies on an external evaluator.
Reported gains under search-aware evaluation
Systems that generate many candidates and report the best one need search-aware evaluation. FINSABER finds previously reported advantages of LLM investing strategies deteriorate over two decades and more than a hundred symbols. Gençay (2026) deflates every reported result by the agent’s own trial count and rejects every LLM-discovered strategy it tests. The same paper shows a deliberately leaky oracle with a Sharpe ratio of 35 passing both the deflated Sharpe ratio and the probability-of-backtest-overfitting test. Statistical correction does not detect look-ahead.
Look-ahead in language models is measurable
Model families with annual training cutoffs, such as DatedGPT and ChronoGPT, let a backtest use only what a model could have known. A recall-based test estimates the probability that a model has internalised a realised outcome. The effect is not always the expected one: Glasserman and Lin find that in-sample, general knowledge of a company distorts sentiment more than specific knowledge of its later returns.
Further reading
- Russell & Wefald (1991), Principles of Metareasoning: the value of computation as a decision-theoretic quantity, and the case for spending computation where it changes the action.
- Hay, Russell, Tolpin & Shimony (2012), Selecting Computations: metalevel control as a Bayesian selection problem, with the argument that the bandit framing is the wrong one.
- Callaway et al. (2017), Learning to Select Computations: a learned approximation to optimal metareasoning, bracketed by the myopic value of information and the value of perfect information.
- Cotton (2026), Wavefront Pruning in Budgeted Brownian Races: the value of computation for many parallel paths under a path-time budget, solved in the mean-field limit.
- Cotton (2026), Consistent Objectives for Self-Improving Systems: which per-step objectives select the same trajectories as the long-term goal, comparing absolute, relative and evolutionary assessment.
- Turtel et al. (2026), Future-as-Label: resolved outcomes as a reward, with the information cutoff enforced on the predictor.
- Li et al. (2025), R&D-Agent-Quant: a complete research loop over factors and models, and the most direct template for an automated quantitative researcher.
- Gençay (2026), What Survives Honest Evaluation?: what such a loop looks like after search and leakage are accounted for.
- Ren et al. (2026), Self-Improvements in Modern Agentic Systems: a survey that organises adaptation by what is updated, from model weights to prompts, memory, tools and control logic.