Autonomous Research Loops

Agents that propose, implement, test and interpret, and then decide what to try next.

A research loop has four stages: a hypothesis, an executable implementation, a measured result, and a choice of the next hypothesis. The first three have improved quickly. The fourth, the metalevel choice, is usually a tree search or a bandit over a benchmark score.

General research agents

The AI Scientist (Lu et al. 2024) generates ideas, writes code, runs experiments, writes a paper and reviews it with an automated reviewer, at under $15 per paper. Its successor, The AI Scientist-v2 (Yamada et al. 2025), drops the human-written code templates and adds an agentic tree search managed by an experiment-manager agent. One of three fully generated manuscripts submitted to an ICLR workshop scored above the average human acceptance threshold.

An independent evaluation is less favourable. Beel, Kan and Baumgart (2025) report that 42% of the original system’s experiments failed through coding errors, that its novelty assessments misclassified established ideas as new, and that some generated papers contained hallucinated numerical results.

AIDE (Jiang et al. 2025) frames machine-learning engineering as optimisation over code, and trial and error as tree search over candidate solutions that reuses and refines the promising ones. On MLE-bench, 75 Kaggle competitions with human leaderboards, the best configuration at publication was o1-preview with AIDE scaffolding, at bronze-medal level or better in 16.9% of competitions.

MLAgentBench (Huang et al. 2023) shows how much the task matters. Success rates run from 100% on established datasets to 0% on Kaggle challenges created after the underlying model was trained. ScienceFlow (Zhao et al. 2026) addresses long-horizon continuity, recovery from dead ends and value-driven compute allocation, and reports 70.22% any-medal on MLE-bench within 24 hours.

Evolutionary program search

FunSearch (Romera-Paredes et al. 2023) pairs a language model with an evaluator and searches over programs rather than answers. AlphaEvolve (Novikov et al. 2025) extends the approach to an evolutionary pipeline that edits code directly, with feedback from one or more evaluators. It found a procedure for multiplying two $4 \times 4$ complex-valued matrices with 48 scalar multiplications.

The same loop can target learning algorithms. DiscoPOP (Lu et al. 2024) has a model propose and implement preference-optimisation losses from the scores of earlier proposals, and finds a loss that blends logistic and exponential terms.

Program search works where the evaluator is cheap, exact and hard to fool. A research loop over noisy financial or economic data has none of those properties, which is why the evaluation literature matters more there.

Quantitative research agents

Qlib (Yang et al. 2020) supplies the infrastructure: data handling, model training and backtesting for AI-driven quantitative research.

RD-Agent(Q) (Li et al. 2025) runs the full loop on top of it. A research stage sets goals, forms hypotheses and maps them to tasks. A development stage writes and runs the code in backtests. A feedback stage evaluates the results and a multi-armed bandit scheduler picks the next direction. The paper reports up to twice the annualised return of classical factor libraries using 70% fewer factors.

Its own limitations section lists the gaps. The framework “relies solely on the LLM’s internal financial knowledge,” and the authors name alternative data such as news sentiment, macroeconomic indicators and corporate filings, domain knowledge, and online adaptation as future work.

AlphaAgent (Tang et al. 2025) constrains factor discovery to resist decay. It enforces originality by abstract-syntax-tree similarity against existing factors, checks consistency between the stated economic hypothesis and the generated factor, and limits structural complexity. EVOQUANT (Mao et al. 2026) edits existing strategies under a multi-stage verification pipeline and distils what it learns into reusable knowledge. It reports the average test Sharpe ratio across seven strategies rising from −0.298 to 0.538.

Trading agents

A separate line builds agents that trade rather than research. FinMem adds layered memory, FinAgent multimodal inputs with dual-level reflection, FinCon a manager-analyst hierarchy with verbal reinforcement, and TradingAgents a simulated trading firm of analysts, bull and bear researchers and a risk team. Each reports gains over baselines. They are useful as references for memory, retrieval and reflection designs.

Their evaluations are the weak point. Nguyen and Pham (2026) survey twelve multi-agent trading systems and document five pervasive evaluation failures: look-ahead bias, survivorship bias, backtest overfitting, neglect of transaction costs and blindness to regime shifts. They show these can reverse the sign of reported returns. FINSABER (Li et al. 2025) re-tests timing strategies over two decades and more than a hundred symbols, and finds the reported advantages deteriorate.

Outputs of an experiment

A research loop learns only from experiments that return three things: an executable artifact, a statement of the information it needed, and a measured result on an interface the agent does not control. An agent that reports a plausible narrative has not yet produced evidence of learning.