Evaluating a Search

The best of many trials is biased upward, and a model trained on the future cannot be tested on it.

A self-improving research system is a search procedure. It tries many variants, keeps the ones that score well, and adapts to what it has seen. Every such procedure produces an optimistic estimate of its own best result. The unit of evaluation is therefore the whole research-and-selection procedure through time, not the winning variant.

Data snooping and multiple testing

White (2000) introduced the reality check: a bootstrap test of whether the best model found by a specification search truly beats a benchmark, accounting for the search that found it. Hansen (2005) improves its power with the test for superior predictive ability, which reduces the influence of poor and irrelevant alternatives. Romano and Wolf (2005) extend it to a stepwise procedure that identifies which of many strategies beat the benchmark.

Harvey, Liu and Zhu (2016) apply multiple testing to the hundreds of published return factors and argue that a new factor should clear a much higher t-statistic than the conventional 2.0.

Backtest overfitting

Bailey, Borwein, López de Prado and Zhu define the probability of backtest overfitting: the probability that the configuration selected as best in-sample underperforms the median configuration out of sample. They estimate it by combinatorially symmetric cross-validation. Bailey and López de Prado (2014) give the deflated Sharpe ratio, which corrects a reported Sharpe ratio for the number of trials and for non-normal returns.

Both corrections need the trial count. A system that discards its failed experiments cannot compute them.

Adaptive data analysis

Classical inference assumes the hypotheses were fixed before the data were seen. In practice each analysis is chosen after the results of earlier ones. Dwork et al. (2014) show that a holdout can be reused for an exponential number of adaptively chosen queries if its answers are perturbed with techniques from differential privacy. The reusable holdout (Science, 2015) gives the practical version.

The guarantees rest on independent samples from a fixed distribution. Financial time series are neither independent nor stationary, so the methods identify the problem more reliably than they solve it there. Once an agent sees a period’s results and adapts to them, that period is no longer untouched evidence for the adapted system.

Leakage

Kapoor and Narayanan (2022) find data leakage in 17 fields, affecting 329 papers, and give a taxonomy of eight types. In their civil-war prediction case study, every paper claiming that complex models beat logistic regression failed to reproduce once leakage was removed.

Look-ahead inside language models

Restricting retrieved documents by date does not produce a clean historical test when the model itself was trained on the period. Glasserman and Lin (2023) separate two effects in news-sentiment strategies: look-ahead, where the model knows what followed, and distraction, where general knowledge of a company interferes with reading the text. Anonymised headlines outperform in-sample, so distraction was the larger effect in their data.

Several tools measure or remove the bias. ChronoBERT and ChronoGPT (He et al. 2025) are trained only on text available at each point in time, and find the bias modest in a next-day return application. DatedGPT (Yan et al. 2026) is a family of twelve 1.3B models with annual cutoffs from 2013 to 2024. It estimates a look-ahead premium of 26.4 basis points per standard deviation for models whose training covers the outcome period.

Gao, Jiang and Yan (2025) estimate a look-ahead propensity from date-only recall queries and test whether forecast accuracy interacts with it. Look-Ahead-Bench (Benhenda 2026) measures the bias through performance decay across market regimes. FinCAD (Li, Wang & Ma 2026) attenuates memorised outcomes at inference time without retraining.

Search-aware evaluation of research agents

Gençay (2026) builds the corrections into the system. The agent acts only through registry-validated tools whose feature space excludes look-ahead, every evaluation is recorded, and all reported performance is deflated by the trial count. As the search proceeds, the best in-sample Sharpe ratio climbs, and the deflation threshold driven by the agent’s own search climbs faster.

Across a 453-stock point-in-time universe and a 39-ETF universe with realistic costs, the evaluation certifies passive benchmarks and rejects every LLM-discovered strategy, across two frontier models, budgets of up to a hundred candidates and five repeated runs. A deliberately leaky oracle with a Sharpe ratio of 35 passes both the deflated Sharpe ratio and the probability-of-backtest-overfitting test. Structural guards against look-ahead and statistical corrections for search catch different failures, and a credible evaluation needs both.

FINSABER (Li et al. 2025) re-tests LLM timing strategies over two decades and more than a hundred symbols, and finds their reported advantages deteriorate, with strategies too conservative in bull markets and too aggressive in bear markets. Nguyen and Pham (2026) propose minimum evaluation standards for multi-agent trading systems.

Prospective evaluation

The cleanest test is a forecast made before the outcome exists. ForecastBench asks only about unresolved events. MLE-bench studies contamination from pretraining alongside its main results. A research system can run the same design internally. Shadow forecasts are recorded and frozen at the time, and a change is promoted only on later evidence, with gaps between training and test periods wide enough for overlapping labels to resolve.