Self-Improvement Mechanisms
Prompts, memory, code, control logic and weights, and the feedback each one needs.
Ren et al. (2026) describe an agent as a configuration that couples a foundation model with a scaffold of prompts, memory, tools and control logic. Self-improvement is an update operator that commits changes to the model parameters or to scaffold components.
Gao et al. (2025) organise the same field by what, when and how an agent evolves. Tayal et al. (2026) survey the subset driven by natural-language feedback, under the name verbal reinforcement learning.
Prompts and programs
DSPy (Khattab et al. 2023) represents a language-model pipeline as a graph of declarative modules and compiles it against a metric, bootstrapping demonstrations. Compiled pipelines beat standard few-shot prompting by over 25% for GPT-3.5 and 65% for Llama-2-13b-chat in its case studies.
OPRO (Yang et al. 2023) puts previous solutions and their scores into the prompt and asks the model for a better one. Promptbreeder (Fernando et al. 2023) evolves task prompts, and also evolves the mutation prompts that produce them. TextGrad (Yuksekgonul et al. 2024) propagates textual feedback backwards through a computation graph, by analogy with automatic differentiation.
GEPA (Agrawal et al. 2025) samples execution traces, reflects on them in natural language to diagnose failures, proposes and tests prompt updates, and combines lessons from a Pareto frontier of its own attempts. Across six tasks it beats GRPO by 6% on average and by up to 20%, using up to 35 times fewer rollouts. Components with identifiable failure cases, such as extraction, routing and classification steps, suit this kind of reflection.
Memory
Reflexion (Shinn et al. 2023) has an agent write a reflection on each failed attempt and keep it in an episodic buffer for later trials, and reports 91% pass@1 on HumanEval. ExpeL (Zhao et al. 2023) extracts insights across a collection of training tasks and recalls them at inference. Voyager (Wang et al. 2023) stores verified behaviours as a library of executable skills.
MemRL (Zhang et al. 2026) addresses retrieval that matches on semantics and returns noise. It learns, from environmental feedback, which stored episodes are useful, and improves at runtime with the model weights fixed.
For a research system, useful memory is often procedural: a record that one extraction procedure handles a class of document correctly, or that a set of historical analogues improved a class of forecast. Learning of this kind needs no weight update.
Code and agent design
Gödel machines (Schmidhuber 2003) rewrite any part of their own code once they have proved the rewrite useful. Proving that most changes are beneficial is impossible in practice, as Zhang et al. (2025) note.
STOP (Zelikman et al. 2023) runs a language-model scaffold that improves programs on itself. The language model proposes strategies including beam search, genetic algorithms and simulated annealing. The authors note this is not full recursive self-improvement, since the model is unchanged, and measure how often generated code bypasses a sandbox.
Meta Agent Search (Hu, Lu & Clune 2024) has a meta agent program new agents in code from an archive of earlier discoveries. Gödel Agent (Yin et al. 2024) lets an agent modify its own logic under a high-level objective. The Darwin Gödel Machine (Zhang et al. 2025) replaces the proof requirement with empirical validation on coding benchmarks and keeps an open-ended archive of variants. It raises SWE-bench performance from 20.0% to 50.0%.
Weights
Agent Lightning (Luo et al. 2025) decouples agent execution from reinforcement-learning training, so an existing agent can be trained from its trajectories with little code change. Agent Lightning v1.0 (He et al. 2026) lets the deployment harness own the interaction loop, and reports Qwen3.5-9B rising from 41.8% to 56.4% on SWE-bench Verified with 6,000 training examples.
Weights can also be trained on resolved forecasts, as in outcome-based reinforcement learning. Hospedales et al. (2020) survey the older meta-learning literature, which improves a learning algorithm from many learning episodes.
Self-correction without external feedback
Huang et al. (2023) study intrinsic self-correction, in which a model revises its answer using only its own judgement. On reasoning tasks models struggle to improve this way, and their performance sometimes falls.
Every method with a reported gain in this literature uses an external signal: a test suite, a benchmark score, an environment reward or a resolved outcome. For a research system that signal comes from the evaluation service, which is why all proposed changes pass through one interface and why failed changes are kept.