Implementations
Open code for the loops, the update mechanisms, the substrate and the evaluations.
Star counts and last-push dates were read from the GitHub API on 30 September 2026. A repository with no recent push may be complete rather than abandoned.
Research loops
| Project | Purpose | Paper | Stars | Last push |
|---|---|---|---|---|
| R&D-Agent, including RD-Agent(Q) | Hypothesis, code, backtest and feedback loop for data-driven research | 2505.15155 | 14,816 | 2026-09-30 |
| Qlib | Quantitative research platform: data, models, backtests | 2009.11189 | 49,078 | 2026-09-22 |
| AlphaAgent | Regularised LLM factor mining | 2502.16789 | 418 | 2026-07-03 |
| AIDE | Tree search over code for ML engineering | 2502.13138 | 1,550 | 2026-09-03 |
| The AI Scientist | End-to-end automated papers, with template code | 2408.06292 | 14,645 | 2025-12-19 |
| The AI Scientist-v2 | Agentic tree search, no templates | 2504.08066 | 7,247 | 2025-12-19 |
| TradingAgents | Multi-agent simulated trading firm | 2412.20138 | 109,371 | 2026-09-29 |
Self-improvement
| Project | Purpose | Paper | Stars | Last push |
|---|---|---|---|---|
| DSPy | Declarative LM pipelines compiled against a metric | 2310.03714 | 38,437 | 2026-09-30 |
| GEPA | Reflective prompt and code evolution | 2507.19457 | 6,822 | 2026-09-29 |
| TextGrad | Textual feedback through computation graphs | 2406.07496 | 3,750 | 2025-07-25 |
| OPRO | Optimisation by prompting | 2309.03409 | 780 | 2024-12-04 |
| Reflexion | Verbal reflection in episodic memory | 2303.11366 | 3,290 | 2025-01-14 |
| MemRL | Reinforcement learning on episodic memory | 2601.03192 | 176 | 2026-07-18 |
| ADAS | Meta Agent Search over agent code | 2408.08435 | 1,638 | 2025-01-28 |
| Gödel Agent | Self-modifying agent | 2410.04444 | 228 | 2025-09-17 |
| Darwin Gödel Machine | Archive-based self-improving coding agents | 2505.22954 | 2,385 | 2025-08-13 |
| Agent Lightning | RL training decoupled from agent execution | 2508.03680 | 18,540 | 2026-09-29 |
Data and supervision
| Project | Purpose | Paper | Stars | Last push |
|---|---|---|---|---|
| RelBench | Relational deep learning benchmark | 2407.20060 | 392 | 2026-09-17 |
| PyTorch Frame | Tabular deep learning used by RDL models | — | 799 | 2026-09-07 |
| Featuretools | Deep Feature Synthesis | — | 7,687 | 2026-09-11 |
| Snorkel | Weak supervision from labelling functions | 1711.10160 | 6,011 | 2026-09-14 |
Decisions and causal estimation
| Project | Purpose | Paper | Stars | Last push |
|---|---|---|---|---|
| SPO | Reference code for SPO and SPO+ | 1710.08005 | 91 | 2024-08-09 |
| Open Bandit Pipeline | Off-policy evaluation estimators and data | — | 711 | 2024-06-03 |
| d3rlpy | Offline RL algorithms including CQL and IQL | — | 1,683 | 2025-09-10 |
| DoubleML | Double/debiased machine learning | — | 791 | 2026-09-23 |
| EconML | Heterogeneous treatment effect estimation | — | 4,803 | 2026-09-28 |
Evaluation
| Project | Purpose | Paper | Stars | Last push |
|---|---|---|---|---|
| MLE-bench | 75 Kaggle competitions for ML agents | 2410.07095 | 1,759 | 2026-04-24 |
| MLAgentBench | 13 ML experimentation tasks | 2310.03302 | 354 | 2024-06-19 |
| ForecastBench | Dynamic forecasting benchmark on unresolved questions | 2409.19839 | 86 | 2026-09-30 |
Reading lists
| Project | Purpose | Paper | Stars | Last push |
|---|---|---|---|---|
| Awesome Self-Improving Agents | Tracks the Ren et al. survey | 2607.13104 | 521 | 2026-09-30 |
Corrections to maintenance status are welcome as an issue.