Learning from Resolved Outcomes
Resolved forecasts supply labels without annotation.
A system that records what it knew at each date, and what it predicted, receives a stream of labelled examples as events resolve. The same stream prices the research that preceded each forecast. A research action that changed a forecast in the right direction was worth something, and one that changed nothing was not.
Proper scoring rules
Brier (1950) introduced the squared-error score for probability forecasts in weather verification. Gneiting and Raftery (2007) give the general theory. A scoring rule is strictly proper when a forecaster maximises expected score only by reporting their true belief. The logarithmic score, the Brier score and the continuous ranked probability score are all strictly proper.
Propriety is what makes a score usable as a training reward. A reward that is not proper can be gamed by a forecaster who reports something other than their belief.
Forecasting benchmarks on unresolved events
Autocast (Zou et al. 2022) pairs questions from forecasting tournaments with a news corpus organised by date, so that a model can be shown only what a human forecaster could have seen. Language models at the time performed far below the human expert baseline.
Halawi et al. (2024) build a retrieval-augmented system that searches for information, forecasts and aggregates. On questions published after the models’ knowledge cutoffs it nears the crowd aggregate of competitive forecasters, and in some settings surpasses it. Schoenegger et al. (2024) find an ensemble of twelve models statistically indistinguishable from a crowd of 925 human forecasters on 31 binary questions.
ForecastBench (Karger et al. 2024) removes leakage by construction. It asks only about events with no known answer at submission time. On its human-comparison subset, expert forecasters outperform the best language model with $p < 0.001$.
Question generation and resolution can themselves be automated. Bosse et al. (2026) use web research agents to write 1,499 questions and resolve them months later. They estimate about 96% of the questions are verifiable and unambiguous, and resolution is about 95% accurate.
Training on outcomes
Three papers from one group trace a path from ranking to reinforcement learning. Turtel, Franklin and Schoenegger (2025) have a model generate pairs of reasoning traces and forecasts for questions that resolve after its knowledge cutoff, rank each pair by distance to the outcome, and fine-tune with direct preference optimisation. Accuracy improves by 7–10% over the base model and over a control fine-tuned on randomised labels.
Outcome-based reinforcement learning (Turtel et al. 2025) applies reinforcement learning with verifiable rewards to prediction-market questions, with news headlines as context. A 14B model matches or surpasses frontier models on accuracy and improves calibration substantially. In a trading simulation on the test questions, the authors estimate a return on investment above 10%.
Future-as-Label (Turtel et al. 2026) states the principle in general form. A predictor sees only causally masked information, a resolver uses later information to establish the outcome, and a proper scoring rule supplies the reward. Qwen3-32B trained this way improves its Brier score by 27% and halves its calibration error relative to the pretrained model. It outperforms Qwen3-235B on constructed future-event tasks and on a Metaculus benchmark.
The paper lists its limitations. Training is offline on resolved events, deployment-time feedback loops “are not evaluated in this study,” and the experiments cover binary outcomes only. Richer outcome spaces and fully online settings are left for future work.
Delayed and shifting feedback
Outcomes arrive at different horizons, so a system that updates weights on its experts or models works with delayed feedback. Joulani, György and Szepesvári (2013) show that delay increases regret multiplicatively in adversarial problems and additively in stochastic ones, and give black-box reductions from the undelayed case. Cesa-Bianchi and Lugosi (2006) is the standard reference for combining expert predictions online.
Adaptive conformal inference (Gibbs & Candès 2021) maintains the long-run coverage of prediction sets under arbitrary distribution shift by re-estimating a single parameter online. Coverage is a frequency guarantee. It is distinct from conditional accuracy, and from the usefulness of the set for any decision.
Evaluation hazards specific to forecasting
Paleka, Goel, Geiping and Tramèr (2025) argue that claims of human-level forecasting deserve caution. They identify many forms of temporal leakage in evaluation, and a gap between benchmark performance and real-world forecasting. Prospective designs such as ForecastBench answer the first concern. The second remains open.
A date filter on retrieved documents does not remove look-ahead bias inside the model itself, and tools for measuring that bias exist.