Bibliography
Work on choosing what to compute and learn next, and on making the result of that choice believable.
arXiv identifiers were resolved against the arXiv API and journal DOIs against Crossref in September 2026. Venues are given where the arXiv record states them.
The core
- Russell, S., and Wefald, E. (1991). Principles of Metareasoning. Artificial Intelligence 49(1–3), 361–395. Computations are actions valued by their effect on later actions. The rational agent performs the computation of highest net value and acts when none remains positive.
- Hay, N., Russell, S., Tolpin, D., and Shimony, S. E. (2012). Selecting Computations: Theory and Applications. arXiv:1207.5879; UAI 2012. Metalevel control as a Bayesian selection problem rather than a bandit problem, since only the final choice collects reward. Finite sampling bounds in some cases, and a counterexample to the conjecture that an optimal policy always reaches a decision.
- Callaway, F., Gul, S., Krueger, P. M., Griffiths, T. L., and Lieder, F. (2017). Learning to select computations. arXiv:1711.06892. Bayesian metalevel policy search. The value of information lies between the myopic value and the value of perfect information, and that bracket makes a learned metalevel policy tractable.
- Turtel, B., Wilczewski, P., Franklin, D., and Skothiem, K. (2026). Future-as-Label: Scalable Supervision from Real-World Outcomes. arXiv:2601.06336. Resolved outcomes as reward. The predictor sees causally masked information, a resolver establishes the outcome, and a proper scoring rule scores the forecast. A 27% Brier improvement for Qwen3-32B. Offline, binary outcomes only. Deployment-time loops are not evaluated.
- Li, Y., Yang, X., Yang, X., Xu, M., Wang, X., et al. (2025). R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization. arXiv:2505.15155; NeurIPS 2025. A full research loop over factors and models: research, development and feedback stages, with a multi-armed bandit choosing the next direction. Up to twice the annualised return of classical factor libraries with 70% fewer factors. Its limitations section names alternative data, domain knowledge and online adaptation as open.
- Gençay, E. (2026). What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery. arXiv:2608.27734. Leakage-safe tools by construction and deflation by the agent’s own trial count. Rejects every LLM-discovered strategy tested, and shows a leaky oracle with Sharpe 35 passing both the deflated Sharpe ratio and the backtest-overfitting test.
Budgeted races
- Cotton, P. (2026). Wavefront Pruning in Budgeted Brownian Races. Working draft; software at microprediction/brownianbandit. A controller pays per unit time for every Brownian path kept alive and prunes to maximise the expected terminal maximum. A path is kept while its future pivotal value exceeds its carrying cost. The value of computation for parallel rollouts and agent fleets.
- Cotton, P. (2026). Posterior-Predictive Pass-at-k. Note. Pass@k is the expected-maximum payoff of a best-of-k batch. A posterior expectation replaces the downward-biased plug-in and cuts held-out log loss from 3.37 to 0.48.
- Cotton, P. (2026). When the Grass Is Greener: Three-Shot Search on Exponentiated Gaussian Landscapes. microprediction/browniansearch. The search-side sibling: where to spend a few evaluations of one path, paid at the final point.
- Cotton, P. (2026). Clay Shooting: Equilibrium Effort Solving Outstanding Problems. Working draft. The strategic sibling. A Brownian attempt buys success probability at a convex price, and the record of solved problems implies what an attempt is worth to whoever makes it.
Pricing information and computation
- Howard, R. A. (1966). Information Value Theory. IEEE Transactions on Systems Science and Cybernetics 2(1), 22–26. The value of information as the gain in expected utility from observing before deciding.
- Matheson, J. E. (1968). The Economic Value of Analysis and Computation. IEEE Transactions on Systems Science and Cybernetics 4(3), 325–332. Extends the value of information from observations to analysis. The earliest statement of the value of computation.
- Gittins, J. C. (1979). Bandit Processes and Dynamic Allocation Indices. Journal of the Royal Statistical Society, Series B 41(2), 148–164. Index policies for allocating effort among alternatives, optimal under the bandit-process assumptions.
- Horvitz, E. J. (2013). Reasoning About Beliefs and Actions Under Computational Resource Constraints. arXiv:1304.2759; UAI 1987. Decision-theoretic control of inference under resource limits: weigh the costs and benefits of approximation procedures, and of acting on a partial result.
- Hansen, E. A., and Zilberstein, S. (1996). Monitoring Anytime Algorithms. ACM SIGART Bulletin 7(2), 28–33. When to stop an algorithm that improves with time and can be interrupted.
- Tolpin, D. and Shimony, S. E. (2012). MCTS Based on Simple Regret. arXiv:1207.5536; AAAI 2012. UCT minimises cumulative regret, but search needs simple regret. Optimising the sampling is itself a metareasoning problem.
- Sezener, E. and Dayan, P. (2020). Static and Dynamic Values of Computation in MCTS. arXiv:2002.04335; UAI 2020. Values Monte Carlo simulations by their expected effect on the action eventually chosen, beyond the myopic horizon.
Resource-rational metareasoning
- Lieder, F., and Griffiths, T. L. (2017). Strategy Selection as Rational Metareasoning. Psychological Review 124(6), 762–794. Human strategy choice modelled as learned metareasoning.
- Lieder, F., Shenhav, A., Musslick, S., and Griffiths, T. L. (2018). Rational Metareasoning and the Plasticity of Cognitive Control. PLOS Computational Biology 14(4), e1006043. Cognitive control adapted by learning which computations are worth their cost.
- Lieder, F., and Griffiths, T. L. (2020). Resource-Rational Analysis: Understanding Human Cognition as the Optimal Use of Limited Computational Resources. Behavioral and Brain Sciences 43. The research programme, with open peer commentary.
Information acquisition
- Frazier, P. I., Powell, W. B., and Dayanik, S. (2008). A Knowledge-Gradient Policy for Sequential Information Collection. SIAM Journal on Control and Optimization 47(5), 2410–2439. The one-step expected improvement in the value of the best alternative.
- Russo, D. and Van Roy, B. (2014). Learning to Optimize via Information-Directed Sampling. arXiv:1403.5556. Acts to minimise squared expected regret per unit of information about the optimal action. A regret bound that scales with the entropy of the optimal action.
- Frazier, P. I. (2018). A Tutorial on Bayesian Optimization. arXiv:1807.02811. Expected improvement, entropy search and the knowledge gradient, with multi-fidelity and multi-source variants.
- Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A. (2016). Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization. arXiv:1603.06560. Adaptive resource allocation with early stopping over random configurations, reported an order of magnitude faster than Bayesian optimisation on its benchmarks. The non-Bayesian counterpoint.
- Li, Y. and Oliva, J. B. (2020). Active Feature Acquisition with Generative Surrogate Models. arXiv:2010.02433. Which missing features to pay for at prediction time, with a generative surrogate estimating the information gain of each.
- Wan, B., Liu, M., Grigas, P., and Shen, Z.-J. M. (2026). Decision-Focused Sequential Experimental Design: A Directional Uncertainty-Guided Approach. arXiv:2602.05340. Sequential experimental design aimed at decision loss rather than prediction error. Stops earlier than decision-blind designs.
Metareasoning in language models
- Graves, A. (2016). Adaptive Computation Time for Recurrent Neural Networks. arXiv:1603.08983. A network learns how many internal steps to take. On text it spends more computation on harder transitions.
- Snell, C., Lee, J., Xu, K., and Kumar, A. (2024). Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv:2408.03314. Compute-optimal test-time scaling depends on prompt difficulty. More than four times the efficiency of best-of-N.
- Chen, X., Xu, J., Liang, T., He, Z., Pang, J., et al. (2024). Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs. arXiv:2412.21187. Reasoning models spend heavily on easy problems for little benefit.
- De Sabbata, C. N., Sumers, T. R., AlKhamissi, B., Bosselut, A., and Griffiths, T. L. (2024). Rational Metareasoning for Large Language Models. arXiv:2410.05563. A reward built on the value of computation, trained by expert iteration. 20–37% fewer tokens at unchanged performance.
- Sui, Y., He, Y., Cao, T., Han, S., Chen, Y., and Hooi, B. (2025). Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models. arXiv:2502.19918; ACL 2026. A contextual bandit chooses between continuing, backtracking, switching strategy and restarting.
- Wang, Z., Tsang, I., and Qian, H. (2026). AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning. arXiv:2608.27964. Present confidence is a poor guide to whether more computation helps. Learns the future value of computation instead, cutting completion tokens by 95.99% for a 0.4-point accuracy cost on held-out GSM8K.
- Fang, H. and Wang, B. (2026). Exploit More, Explore Smarter for Budget-Constrained Agentic Search. arXiv:2608.23848. Tree expansion as a value-of-information decision under a small evaluation budget.
Learning from resolved outcomes
- Cotton, P. (2022). Microprediction: Building an Open AI Network. MIT Press. Chapter 5, “Micromanagers” (pp. 81–120), considers agents that solicit, compensate and combine the predictions of other agents.
- Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review 78(1), 1–3. The squared-error score for probability forecasts.
- Gneiting, T., and Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association 102, 359–378. The general theory of proper scoring rules, and the reason they can serve as rewards.
- Zou, A., Xiao, T., Jia, R., Kwon, J., Mazeika, M., et al. (2022). Forecasting Future World Events with Neural Networks. arXiv:2206.15474; NeurIPS 2022. Tournament questions paired with a date-ordered news corpus, so a model sees only what a forecaster could have seen.
- Halawi, D., Zhang, F., Chen, Y.-H., and Steinhardt, J. (2024). Approaching Human-Level Forecasting with Language Models. arXiv:2402.18563. A retrieval-augmented forecasting system that nears the crowd aggregate on post-cutoff questions.
- Schoenegger, P., Tuminauskaite, I., Park, P. S., and Tetlock, P. E. (2024). Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy. arXiv:2402.19379. A twelve-model ensemble is statistically indistinguishable from a 925-person crowd on 31 questions.
- Karger, E., Bastani, H., Chen, Y.-H., Jacobs, Z., Halawi, D., et al. (2024). ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities. arXiv:2409.19839. A dynamic benchmark of questions about unresolved events. Experts beat the best language model, $p < 0.001$.
- Bosse, N. I., Mühlbacher, P., Wildman, J., Phillips, L., and Schwarz, D. (2026). Automating Forecasting Question Generation and Resolution for AI Evaluation. arXiv:2601.22444. Agents write and later resolve forecasting questions. About 96% verifiable and 95% correctly resolved.
- Turtel, B., Franklin, D., and Schoenegger, P. (2025). LLMs Can Teach Themselves to Better Predict the Future. arXiv:2502.05253. Self-play forecasts ranked by distance to the outcome, then DPO. 7–10% accuracy gains over base and randomised-label controls.
- Turtel, B., Franklin, D., Skotheim, K., Hewitt, L., and Schoenegger, P. (2025). Outcome-based Reinforcement Learning to Predict the Future. arXiv:2505.17989. Reinforcement learning with verifiable rewards on prediction-market questions. A 14B model matches frontier accuracy with better calibration.
- Joulani, P., György, A., and Szepesvári, C. (2013). Online Learning under Delayed Feedback. arXiv:1306.0686; ICML 2013. Delay multiplies regret in adversarial problems and adds to it in stochastic ones.
- Cesa-Bianchi, N., and Lugosi, G. (2006). Prediction, Learning, and Games. Cambridge University Press. The reference for online combination of expert predictions.
- Gibbs, I. and Candès, E. (2021). Adaptive Conformal Inference Under Distribution Shift. arXiv:2106.00170. Long-run coverage of prediction sets under arbitrary drift. Coverage is not conditional accuracy.
Autonomous research loops
- Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292. Ideas, code, experiments, paper and automated review, under $15 per paper.
- Beel, J., Kan, M.-Y., and Baumgart, M. (2025). Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?. arXiv:2502.14297. Independent evaluation of the AI Scientist: 42% of experiments failed on coding errors, novelty checks misfired, some papers had hallucinated numbers.
- Yamada, Y., Lange, R. T., Lu, C., Hu, S., Lu, C., et al. (2025). The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search. arXiv:2504.08066. Agentic tree search without human code templates. One of three workshop submissions scored above the average acceptance threshold.
- Jiang, Z., Schmidt, D., Srikanth, D., Xu, D., Kaplan, I., et al. (2025). AIDE: AI-Driven Exploration in the Space of Code. arXiv:2502.13138. Machine-learning engineering as tree search over code.
- Chan, J. S., Chowdhury, N., Jaffe, O., Aung, J., Sherburn, D., et al. (2024). MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. arXiv:2410.07095. 75 Kaggle competitions. o1-preview with AIDE reached bronze or better in 16.9%.
- Huang, Q., Vora, J., Liang, P., and Leskovec, J. (2023). MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation. arXiv:2310.03302. Success from 100% on old datasets to 0% on post-training Kaggle challenges.
- Zhao, M., Dong, J., Xu, K., Hasan, Z., Fan, C., et al. (2026). ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond. arXiv:2608.14354. Long-horizon research with recoverable executable states and value-driven compute allocation. 70.22% any-medal on MLE-bench within 24 hours.
- Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., et al. (2023). Mathematical Discoveries from Program Search with Large Language Models. Nature 625, 468–475. FunSearch. Evolves programs against an evaluator.
- Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P.-S., et al. (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv:2506.13131. Evolutionary code editing with evaluator feedback. A 48-multiplication procedure for $4 \times 4$ complex matrices.
- Lu, C., Holt, S., Fanconi, C., Chan, A. J., Foerster, J., et al. (2024). Discovering Preference Optimization Algorithms with and for Large Language Models. arXiv:2406.08414. A model proposes preference-optimisation losses from the scores of earlier ones.
Quantitative research and trading agents
- Yang, X., Liu, W., Zhou, D., Bian, J., and Liu, T.-Y. (2020). Qlib: An AI-oriented Quantitative Investment Platform. arXiv:2009.11189. Infrastructure for AI-driven quantitative research.
- Tang, Z., Chen, Z., Yang, J., Mai, J., Zheng, Y., et al. (2025). AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay. arXiv:2502.16789. Factor mining with originality, hypothesis alignment and complexity constraints to resist decay.
- Mao, J., Li, C., Li, X., Duan, Q., Yuan, J., et al. (2026). EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading. arXiv:2607.12455. Verified strategy edits distilled into reusable knowledge.
- Yu, Y., Li, H., Chen, Z., Jiang, Y., Li, Y., et al. (2023). FinMem: A Performance-Enhanced LLM Trading Agent with Layered Memory and Character Design. arXiv:2311.13743. Layered memory for a trading agent.
- Zhang, W., Zhao, L., Xia, H., Sun, S., Sun, J., et al. (2024). A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and Generalist. arXiv:2402.18485. FinAgent. Multimodal market intelligence with dual-level reflection.
- Yu, Y., Yao, Z., Li, H., Deng, Z., Cao, Y., et al. (2024). FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making. arXiv:2407.06567. Manager-analyst hierarchy with conceptual verbal reinforcement.
- Xiao, Y., Sun, E., Luo, D., and Wang, W. (2024). TradingAgents: Multi-Agents LLM Financial Trading Framework. arXiv:2412.20138. A simulated trading firm of specialised agents.
- Nguyen, P. and Pham, T. (2026). Toward Reliable Evaluation of LLM-Based Financial Multi-Agent Systems: Taxonomy, Coordination Primacy, and Cost Awareness. arXiv:2603.27539. Five evaluation failures in multi-agent trading systems that can reverse the sign of reported returns.
- Li, W. W., Kim, H., Cucuringu, M., and Ma, T. (2025). Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?. arXiv:2505.07078. FINSABER. Reported LLM timing advantages deteriorate over two decades and 100+ symbols.
Self-improvement mechanisms
- Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., et al. (2023). DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. arXiv:2310.03714. Pipelines as declarative modules compiled against a metric.
- Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., et al. (2023). Large Language Models as Optimizers. arXiv:2309.03409; ICLR 2024. OPRO. Past solutions and scores in the prompt.
- Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rocktäschel, T. (2023). Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv:2309.16797. Evolves task prompts and the mutation prompts that produce them.
- Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., et al. (2024). TextGrad: Automatic "Differentiation" via Text. arXiv:2406.07496. Textual feedback propagated backwards through a computation graph.
- Agrawal, L. A., Tan, S., Soylu, D., Ziems, N., Khare, R., et al. (2025). GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning. arXiv:2507.19457; ICLR 2026. Reflective prompt evolution over execution traces. Beats GRPO by 6% on average with up to 35 times fewer rollouts.
- Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., and Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. Verbal reflections kept in episodic memory across trials.
- Zhao, A., Huang, D., Xu, Q., Lin, M., Liu, Y.-J., and Huang, G. (2023). ExpeL: LLM Agents Are Experiential Learners. arXiv:2308.10144; AAAI 2024. Insights extracted across tasks and recalled at inference.
- Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. A growing library of executable skills.
- Zhang, S., Wang, J., Zhou, R., Liao, J., Feng, Y., et al. (2026). MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory. arXiv:2601.03192. Learns which episodic memories are useful from environmental feedback, with weights fixed.
- Schmidhuber, J. (2003). Gödel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements. arXiv:cs/0309048. Provably optimal self-rewrites, conditional on finding the proof.
- Zelikman, E., Lorch, E., Mackey, L., and Kalai, A. T. (2023). Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. arXiv:2310.02304; COLM 2024. A scaffold that improves itself, with sandbox-bypass rates measured.
- Hu, S., Lu, C., and Clune, J. (2024). Automated Design of Agentic Systems. arXiv:2408.08435. A meta agent programs new agents from an archive of discoveries.
- Yin, X., Wang, X., Pan, L., Lin, L., Wan, X., and Wang, W. Y. (2024). Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement. arXiv:2410.04444; ACL 2025. An agent that modifies its own logic under a high-level objective.
- Zhang, J., Hu, S., Lu, C., Lange, R., and Clune, J. (2025). Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv:2505.22954. Empirical validation replaces proof. SWE-bench 20.0% to 50.0%.
- Luo, X., Zhang, Y., He, Z., Wang, Z., Zhao, S., et al. (2025). Agent Lightning: Train ANY AI Agents with Reinforcement Learning. arXiv:2508.03680. Agent execution decoupled from RL training.
- He, Z., Zhang, S., Zhou, Z., Yang, Y., Kang, Y., et al. (2026). Agent Lightning v1.0: Towards Harnessed Agentic RL. arXiv:2608.17528. The deployment harness owns the loop. SWE-bench Verified 41.8% to 56.4% for Qwen3.5-9B.
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. Without external feedback, models struggle to correct their reasoning and sometimes get worse.
- Hospedales, T., Antoniou, A., Micaelli, P., and Storkey, A. (2020). Meta-Learning in Neural Networks: A Survey. arXiv:2004.05439. The meta-learning literature before language agents.
Relational features and weak labels
- Kanter, J. M., and Veeramachaneni, K. (2015). Deep Feature Synthesis: Towards Automating Data Science Endeavors. IEEE DSAA 2015. Stacked aggregations across related tables. The explicit baseline.
- Fey, M., Hu, W., Huang, K., Lenssen, J. E., Ranjan, R., et al. (2023). Relational Deep Learning: Graph Representation Learning on Relational Databases. arXiv:2312.04615. A database as a temporal heterogeneous graph, one node per row.
- Robinson, J., Ranjan, R., Hu, W., Huang, K., Han, J., et al. (2024). RelBench: A Benchmark for Deep Learning on Relational Databases. arXiv:2407.20060. The benchmark, with a user study against a data scientist’s hand-built features.
- Gu, J., Ranjan, R., Kanatsoulis, C., Tang, H., Jurkovic, M., et al. (2026). RelBench v2: A Large-Scale Benchmark and Repository for Relational Data. arXiv:2602.12606; ICLR 2026. 11 datasets, 22 million rows, autocomplete tasks under temporal constraints.
- Zhang, Y., Xu, L., Gan, Q., Wipf, D., and Wang, M. (2026). RDBLearn: Simple In-Context Prediction Over Relational Databases. arXiv:2602.18495. Relational aggregations plus a tabular foundation model, competitive with supervised baselines.
- Gentzkow, M., Kelly, B., and Taddy, M. (2019). Text as Data. Journal of Economic Literature 57(3), 535–574. Statistical methods for text in economics.
- Ke, Z. T., Kelly, B., and Xiu, D. (2019). Predicting Returns with Text Data. NBER Working Paper 26186. Sentiment learned from subsequent returns rather than from a dictionary.
- Lopez-Lira, A. and Tang, Y. (2023). Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models. arXiv:2304.07619. GPT-4 scores capture initial reactions to post-cutoff headlines. Returns decline as adoption rises.
- Ratner, A., Bach, S. H., Ehrenberg, H., Fries, J., Wu, S., and Ré, C. (2017). Snorkel: Rapid Training Data Creation with Weak Supervision. arXiv:1711.10160. Labelling functions of unknown accuracy, denoised without ground truth.
From forecasts to decisions
- Elmachtoub, A. N., and Grigas, P. (2022). Smart “Predict, then Optimize”. Management Science 68(1), 9–26; arXiv:1710.08005. Prediction models trained on decision error rather than prediction error. SPO+ is a convex, statistically consistent surrogate.
- Gârleanu, N., and Pedersen, L. H. (2013). Dynamic Trading with Predictable Returns and Transaction Costs. Journal of Finance 68(6), 2309–2340. The optimal portfolio trades partway toward an aim portfolio, and persistent signals get more weight.
- Dudik, M., Langford, J., and Li, L. (2011). Doubly Robust Policy Evaluation and Learning. arXiv:1103.4601; ICML 2011. Accurate policy values if either the reward model or the logging-policy model is good.
- Swaminathan, A. and Joachims, T. (2015). Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. arXiv:1502.02362. Propensity-weighted learning with a variance penalty.
- Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020). Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. arXiv:2005.01643. The tutorial on learning policies from fixed data.
- Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020). Conservative Q-Learning for Offline Reinforcement Learning. arXiv:2006.04779. A value function that lower-bounds the true value.
- Kostrikov, I., Nair, A., and Levine, S. (2021). Offline Reinforcement Learning with Implicit Q-Learning. arXiv:2110.06169. Never queries actions outside the dataset.
- Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/Debiased Machine Learning for Treatment and Structural Parameters. The Econometrics Journal 21(1), C1–C68. Orthogonal scores and cross-fitting. Estimation under stated identification assumptions, not a substitute for them.
Evaluating a search
- White, H. (2000). A Reality Check for Data Snooping. Econometrica 68(5), 1097–1126. Tests whether the best model from a search beats a benchmark, accounting for the search.
- Hansen, P. R. (2005). A Test for Superior Predictive Ability. Journal of Business & Economic Statistics 23(4), 365–380. More power than the reality check, less sensitive to poor alternatives.
- Romano, J. P., and Wolf, M. (2005). Stepwise Multiple Testing as Formalized Data Snooping. Econometrica 73(4), 1237–1282. Which of many strategies beat the benchmark.
- Harvey, C. R., Liu, Y., and Zhu, H. (2016). … and the Cross-Section of Expected Returns. Review of Financial Studies 29(1), 5–68. Multiple testing applied to hundreds of published factors.
- Bailey, D. H., and López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management 40(5), 94–107. A Sharpe ratio corrected for the number of trials.
- Bailey, D. H., Borwein, J., López de Prado, M., and Zhu, Q. J. (2016). The Probability of Backtest Overfitting. Journal of Computational Finance. The chance that the in-sample winner is below median out of sample.
- Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A. (2014). Preserving Statistical Validity in Adaptive Data Analysis. arXiv:1411.2664. Holdout reuse for adaptively chosen queries, via differential privacy.
- Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A. (2015). The Reusable Holdout: Preserving Validity in Adaptive Data Analysis. Science 349(6248), 636–638. The practical version. Assumes independent samples from a fixed distribution.
- Kapoor, S. and Narayanan, A. (2022). Leakage and the Reproducibility Crisis in ML-based Science. arXiv:2207.07048. Leakage in 17 fields and 329 papers, and a taxonomy of eight types.
- Paleka, D., Goel, S., Geiping, J., and Tramèr, F. (2025). Pitfalls in Evaluating Language Model Forecasters. arXiv:2506.00723. Temporal leakage and weak extrapolation in evaluations of LLM forecasters.
- Glasserman, P. and Lin, C. (2023). Assessing Look-Ahead Bias in Stock Return Predictions Generated By GPT Sentiment Analysis. arXiv:2309.17322. Distraction by general company knowledge outweighs look-ahead in-sample.
- He, S., Lv, L., Manela, A., and Wu, J. (2025). Chronologically Consistent Large Language Models. arXiv:2502.21206. ChronoBERT and ChronoGPT, trained only on text available at each date.
- Yan, Y., Tang, R., Gao, Z., Jiang, W., and Lu, Y. (2026). DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining. arXiv:2603.11838. Twelve 1.3B models with annual cutoffs. A look-ahead premium of 26.4 basis points per standard deviation.
- Gao, Z., Jiang, W., and Yan, Y. (2025). Detecting Lookahead Bias in LLM Forecasts. arXiv:2512.23847. A look-ahead propensity estimated from date-only recall queries.
- Benhenda, M. (2026). Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance. arXiv:2601.13770. Look-ahead measured by performance decay across regimes.
- Li, W. W., Wang, M., and Ma, T. (2026). Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models. arXiv:2605.24564; EMNLP 2026. FinCAD. Attenuates memorised outcomes at inference.
Surveys
- Gao, H., Geng, J., Hua, W., Hu, M., Juan, X., et al. (2025). A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence. arXiv:2507.21046; TMLR 2026. What, when and how agents evolve.
- Ren, Z., Chen, Y., Guo, D., Rong, G., Li, T., et al. (2026). Self-Improvements in Modern Agentic Systems: A Survey. arXiv:2607.13104. Self-improvement as an update operator on model parameters or scaffold components.
- Tayal, K., Sharma, A., Winata, G. I., Das, A., and Sahu, S. (2026). The Rise of Verbal Reinforcement Learning. arXiv:2609.01597. Natural-language feedback as grounding, deliberation and learning signal.