Introduction
Two levels of reasoning, the value of a computation, and the neighbouring vocabularies.
Object level and meta level
An object-level question is about the world: what a filing implies for revenue, whether a policy change will raise demand, how likely an event is to occur by a given date. A meta-level question is about the system’s own effort: which of the available investigations to run next, and whether to run any at all.
A research system faces the meta-level question constantly. It can read another source, check an inventory figure, fit a model, look for a disconfirming report, revisit a historical analogue, or stop. Each of these costs time or money. Each is worth doing only if it might change what the system does next.
The value of computation
Let $s$ be the system’s current information state and $U(a)$ the utility of action $a$. Acting now yields $\max_a \mathbb{E}[U(a) \mid s]$. A computation $c$ produces a result that updates the state to $s \cdot c$, and the system then acts on the updated state. The value of the computation is the expected gain from acting after it, less its cost:
The first term averages over the results the computation might return. A computation whose every possible result leaves the chosen action unchanged has zero gross value, however much uncertainty it removes. Reducing uncertainty about something irrelevant to the decision earns nothing.
Howard (1966) gave the same expression for the value of information, and Matheson (1968) applied it to analysis and computation. Russell and Wefald (1991) made it the basis of a general theory of metareasoning, in which an agent performs the computation of highest value and acts when none has positive value.
Myopia and intractability
The rule is simple to state and hard to apply. The value of one computation depends on which computations follow it, so the exact problem is a sequential decision over the whole space of computation sequences. The usual approximation is myopic. It values each computation as if the system will act immediately after it.
Myopic estimates undervalue computations that pay off only in combination, and a myopic policy can stop too early. Callaway et al. (2017) note that the true value of information lies between the myopic value and the value of perfect information, and use that bracket to learn a metalevel policy. Hay et al. (2012) give a counterexample to the intuitive conjecture that an optimal metalevel policy always reaches a decision.
Learned metareasoning
Exact metareasoning is intractable, so the practical question is whether a metalevel policy can be learned from experience. The ingredients are a log of past computations, the information state each was run in, and a later outcome that says whether the resulting decision was good.
Several recent systems learn such a policy for a language model’s own reasoning. De Sabbata et al. (2024) train a model with a reward that penalises reasoning whose value of computation does not cover its cost. AERA (2026) learns whether generating another block of candidate answers is likely to recover a better one, using future correctness only to build offline supervision.
A research system is the same problem at a larger grain. The computations are experiments, queries and model fits rather than reasoning tokens. The outcomes arrive weeks or months later, and the resolved-outcomes literature shows how to turn them into training signal.
Neighbouring vocabularies
| Term | Emphasis | Representative work |
|---|---|---|
| Rational metareasoning | Choosing computations by their expected value, and deciding when to stop | Russell & Wefald 1991; Hay et al. 2012 |
| Resource-rational analysis | Cognition as the optimal use of limited computation | Lieder & Griffiths 2020 |
| Autonomous research | Generating hypotheses, running experiments and evaluating the results | Lu et al. 2024; Li et al. 2025 |
| Self-improving agents | Updating prompts, memory, tools, procedures or weights from experience | Gao et al. 2025; Ren et al. 2026 |
| Meta-learning | Using many learning episodes to improve the learning algorithm itself | Hospedales et al. 2020 |
| Online and continual learning | Incorporating evidence as it arrives, including delayed feedback and drift | Joulani et al. 2013; Gibbs & Candès 2021 |
“Self-improving quantitative research” names an application: a system that learns both how to predict and how to conduct the research that produces its predictions. “Learned metareasoning” names the capability that distinguishes it, a research policy that improves its own allocation of effort. The central question of the subject is which next research action has the greatest expected value, given what is currently known.
A reference architecture
The controller never touches data or scores directly. Every proposal passes through one interface.
The separation has precedents. Future-as-Label gives the predictor only information before a cutoff and lets a separate resolver use later information. RD-Agent(Q) divides its loop into research, development and feedback stages. Gençay (2026) lets the agent act only through registry-validated tools whose feature space excludes look-ahead, and records every evaluation the search performs.
Separating a research controller from the evaluation of its proposals lets the system learn from many more observations than completed decisions. Every resolved forecast scores the research that fed it, whether or not any action followed.
Constraints on the data
The literature on evaluating a search implies a short list of requirements for any system that learns from its own history.
- Reproducible information states. A record of when something happened, when the system first received it, and which version was visible at each time. A daily snapshot cannot recover an intraday sequence that was never recorded. The Autocast news corpus is organised by date for this reason.
- Tasks defined before outcomes are seen. The task set includes items that were ignored and opportunities that produced no action, so that it is not built around events that later proved dramatic. Paleka et al. (2025) catalogue the forms of temporal leakage that follow otherwise.
- Explicit targets. A probability, a duration or a magnitude has a denominator, a horizon and an uncertainty, and is scored with a proper scoring rule.
- Clustering of repeated coverage. Many reports of one event are one observation, not many independent successes.
- Retention of failed experiments. The number of trials sets the bar a reported result must clear, as in the deflated Sharpe ratio.
- Consumed holdouts. Once a system adapts to a period’s results, that period is no longer untouched evidence for the adapted system (Dwork et al. 2015).