From Forecasts to Decisions
Equal forecast errors can have unequal costs, and logged decisions reflect the policy that made them.
Decision-focused learning
The usual pipeline predicts and then optimises, and trains the predictor to minimise prediction error. Smart “Predict, then Optimize” (Elmachtoub & Grigas; Management Science 2022) trains it on the decision error the prediction induces instead. The SPO loss is hard to optimise, so the paper derives a convex surrogate, SPO+, and proves it statistically consistent with the SPO loss under mild conditions.
The gains are largest when the prediction model is misspecified. In the paper’s shortest-path and portfolio experiments, linear models trained with SPO+ tend to beat random forests even when the ground truth is highly nonlinear.
The same reasoning applies one level up, to the choice of what to learn. Wan et al. (2026) design sequential experiments to reduce decision loss rather than prediction error, and reach a stopping point earlier than decision-blind designs.
Signals with different horizons
Gârleanu and Pedersen (2013) solve for optimal dynamic trading when returns are predictable from signals that decay at different rates and trading is costly. The optimal portfolio trades partway toward an aim portfolio, a weighted average of the current and expected future Markowitz portfolios. Persistent signals receive more weight than fast-decaying ones.
The result gives forecasts at different horizons a common use. A signal that matters over days and one that matters over months enter the same portfolio through their decay rates, so both horizons need calibrated forecasts. Keeping forecast scores separate from portfolio results distinguishes an improvement in forecasting from a change in risk taken.
Learning from logged decisions
Historical feedback reflects the actions an earlier policy chose. Evaluating a new policy on that data means correcting for the mismatch. Doubly robust estimation (Dudík, Langford & Li 2011) combines a model of rewards with a model of the logging policy, and stays accurate if either one is good.
Counterfactual risk minimisation (Swaminathan & Joachims 2015) learns from logged bandit feedback with propensity weighting, and penalises the variance of the weighted estimate through generalisation bounds.
Both depend on coverage. Logged action probabilities make the correction possible where they exist. Estimating them does not recover actions the old policy never took, and does not remove confounding the logs do not record.
Offline reinforcement learning
When decisions change later opportunities, through inventory, execution or information gathered, the problem becomes sequential. Levine, Kumar, Tucker and Fu (2020) review learning policies from fixed datasets, and the central difficulty of distribution shift between the data and the learned policy.
Conservative Q-learning (Kumar et al. 2020) learns a value function that lower-bounds the true value of the policy, so actions poorly supported by the data are not overvalued. Implicit Q-learning (Kostrikov, Nair & Levine 2021) never evaluates actions outside the dataset. It estimates the value of the best in-sample action with an upper expectile.
Observability under price taking
For a price-taking system, later market prices score a directional forecast whether or not a trade followed. That makes many more observations available than completed trades. Fills, market impact and the results of research never performed are not observed, and are the quantities off-policy methods are needed for.
A correctly forecast event may already be reflected in prices. A forecast of the event and a forecast of the market’s surprise at it are different targets and need separate scores.
Causal estimation
Double/debiased machine learning (Chernozhukov et al. 2018) estimates treatment effects with flexible machine-learning models for the nuisance functions, using orthogonal scores and cross-fitting to remove the bias the flexible fits introduce. It estimates effects under stated identification assumptions. It cannot create identification from text and correlations where none exists.