Literature Map

Each node is a paper, coloured by literature. The circle holds the metareasoning literature proper: the classical theory of selecting computations and its descendants inside language models. Edge labels name the relation. Six edges join works that the literature has not yet connected, searched for and not found as of 30 September 2026.

Corrections welcome: file an issue on microprediction/metareasoning.

Gaps

Each gap rests on keyword searches of arXiv titles and abstracts run on 30 September 2026, plus the reference lists of the papers at either end. A gap is withdrawn when the missing work turns up.

  1. Research controllers do not value an experiment by its effect on a decision. The classical theory values a computation by the change it makes to a later action (Russell & Wefald, Hay et al.). Research agents choose directions by tree search or a bandit over a benchmark score (RD-Agent(Q), AIDE). The nearest work, ExTS, treats tree expansion as a value-of-information decision, and ScienceFlow allocates compute by validated progress. Both value against the score being optimised.

    Decision-focused experimental design values experiments by decision loss, but for linear optimisation rather than an agent. The budgeted Brownian race prices each parallel path by its pivotal value for the terminal payoff, as a mean-field theory rather than inside an agent. Confidence: moderate.

  2. No research loop uses an outcome-trained forecaster as its object level. Language models trained on resolved outcomes with proper scoring rules (Future-as-Label, outcome-based RL) and autonomous loops that propose and test hypotheses have not been combined. Searches pairing research agents with Brier scores or resolved outcomes returned only automated question generation, which builds the evaluation rather than the loop. Confidence: moderate.
  3. No off-policy evaluation of research policies. Logged experiments are exactly the data off-policy estimators such as doubly robust evaluation are built for: a logging policy chose which experiment to run, and only its outcome was observed. No paper found evaluates a candidate research controller this way. The nearest hit runs the other way: code-modifying language-model agents that optimise off-policy evaluation itself. Confidence: moderate.
  4. Adaptive-data-analysis methods have not reached research agents. The reusable holdout addresses the exact problem an adaptive agent creates. Searches combining adaptive data analysis or holdout reuse with language models or agents found nothing relevant. Gençay (2026) corrects by deflating for trial count instead. Confidence: moderate to high.
  5. Research agents do not learn from their own experiment logs as relational data. An agent’s history of proposals, runs and outcomes is a linked, timestamped database of the kind relational deep learning targets. RDBLearn ships an agent-facing interface for relational prediction, so the tooling exists. No paper found applies it to an agent’s own logs to predict which proposals will pay off. Confidence: low to moderate.
  6. Search in language-model agents is not decision-focused. Smart Predict-then-Optimize shows that training on decision error beats training on prediction error when models are misspecified. Agentic search methods score candidates by a task metric. None found scores a candidate forecast or feature by the decision loss it induces downstream. Confidence: moderate.