Relational Features and Weak Labels
Linked tables, text, and noisy sources of supervision.
A research system’s history is usually stored in a relational database: documents, events, extracted facts, forecasts, comments, decisions and outcomes, joined by keys and stamped with times. Features have to be built from those linked, timestamped tables, and the many imperfect judgements they hold have to be turned into labels.
Relational databases as temporal graphs
Deep Feature Synthesis (Kanter & Veeramachaneni 2015) builds features automatically by stacking aggregations along the relationships between tables. It is the explicit, interpretable baseline.
Relational Deep Learning (Fey et al. 2023) treats a database as a temporal heterogeneous graph, with a node for each row and an edge for each primary–foreign-key link, and learns across it with message-passing networks. RelBench (Robinson et al. 2024) supplies the benchmark. In its user study, a relational model beat features engineered by an experienced data scientist while cutting the human work by more than an order of magnitude.
RelBench v2 (Gu et al. 2026) grows to 11 datasets, over 22 million rows and 29 tables. It adds autocomplete tasks, which infer missing attributes while respecting temporal constraints, and reports relational models consistently beating single-table baselines.
Explicit aggregation remains competitive. RDBLearn (Zhang et al. 2026) featurises each target row with relational aggregations over linked records, then runs an off-the-shelf tabular foundation model on the result. It is the best foundation-model approach the authors evaluate on RelBench and 4DBInfer, and at times beats supervised baselines trained on each dataset.
For a research database, the comparison that matters is between a learned relational representation and timestamped SQL aggregates fed to a regularised model. Candidate aggregates include disagreement among sources, revisions to an estimate, counts of independent sources, and changes in a recommendation. A graph database is not required for either.
Text as predictive data
Gentzkow, Kelly and Taddy (2019) survey the statistical methods for using text as data in economics: representing documents, and linking them to outcomes through dictionaries, generative models and regression. Ke, Kelly and Xiu (2019) learn a sentiment score from the returns that follow news, rather than from a generic dictionary.
Lopez-Lira and Tang (2023) find that GPT-4 scores on post-cutoff headlines capture initial market reactions, with about 90% portfolio-day hit rates for the non-tradable initial reaction, and predict some subsequent drift. Strategy returns decline as adoption of such models rises. Task-specific signal learned from text is a different object from generic sentiment, and the two should be compared directly.
Weak supervision
Snorkel (Ratner et al. 2017) trains models from labelling functions, heuristics whose accuracies and correlations are unknown, and learns to denoise their outputs without ground truth.
A research system accumulates many such sources: automated recommendations, analyst comments, extraction heuristics and model scores. The Snorkel framing treats them as noisy labellers with unknown and possibly correlated errors, rather than taking any one of them as truth. A source’s estimate stays a prediction until an outcome checks it.