Implementations

Open code for the loops, the update mechanisms, the substrate and the evaluations.

Star counts and last-push dates were read from the GitHub API on 30 September 2026. A repository with no recent push may be complete rather than abandoned.

Research loops

ProjectPurposePaperStarsLast push
R&D-Agent, including RD-Agent(Q)Hypothesis, code, backtest and feedback loop for data-driven research2505.1515514,8162026-09-30
QlibQuantitative research platform: data, models, backtests2009.1118949,0782026-09-22
AlphaAgentRegularised LLM factor mining2502.167894182026-07-03
AIDETree search over code for ML engineering2502.131381,5502026-09-03
The AI ScientistEnd-to-end automated papers, with template code2408.0629214,6452025-12-19
The AI Scientist-v2Agentic tree search, no templates2504.080667,2472025-12-19
TradingAgentsMulti-agent simulated trading firm2412.20138109,3712026-09-29

Self-improvement

ProjectPurposePaperStarsLast push
DSPyDeclarative LM pipelines compiled against a metric2310.0371438,4372026-09-30
GEPAReflective prompt and code evolution2507.194576,8222026-09-29
TextGradTextual feedback through computation graphs2406.074963,7502025-07-25
OPROOptimisation by prompting2309.034097802024-12-04
ReflexionVerbal reflection in episodic memory2303.113663,2902025-01-14
MemRLReinforcement learning on episodic memory2601.031921762026-07-18
ADASMeta Agent Search over agent code2408.084351,6382025-01-28
Gödel AgentSelf-modifying agent2410.044442282025-09-17
Darwin Gödel MachineArchive-based self-improving coding agents2505.229542,3852025-08-13
Agent LightningRL training decoupled from agent execution2508.0368018,5402026-09-29

Data and supervision

ProjectPurposePaperStarsLast push
RelBenchRelational deep learning benchmark2407.200603922026-09-17
PyTorch FrameTabular deep learning used by RDL models—7992026-09-07
FeaturetoolsDeep Feature Synthesis—7,6872026-09-11
SnorkelWeak supervision from labelling functions1711.101606,0112026-09-14

Decisions and causal estimation

ProjectPurposePaperStarsLast push
SPOReference code for SPO and SPO+1710.08005912024-08-09
Open Bandit PipelineOff-policy evaluation estimators and data—7112024-06-03
d3rlpyOffline RL algorithms including CQL and IQL—1,6832025-09-10
DoubleMLDouble/debiased machine learning—7912026-09-23
EconMLHeterogeneous treatment effect estimation—4,8032026-09-28

Evaluation

ProjectPurposePaperStarsLast push
MLE-bench75 Kaggle competitions for ML agents2410.070951,7592026-04-24
MLAgentBench13 ML experimentation tasks2310.033023542024-06-19
ForecastBenchDynamic forecasting benchmark on unresolved questions2409.19839862026-09-30

Reading lists

ProjectPurposePaperStarsLast push
Awesome Self-Improving AgentsTracks the Ren et al. survey2607.131045212026-09-30

Corrections to maintenance status are welcome as an issue.