The AiEdge Newsletter

The AiEdge Newsletter

How Self-Improving Agents Learn from Experience

How executable tasks and adaptive curricula turn production failures into useful practice.

Damien Benveniste's avatar
Damien Benveniste
Oct 09, 2026
∙ Paid

A self-improving agent uses evidence from earlier tasks to change something retained for later tasks, then demonstrates better behavior on new runs. The change might be a prompt, a tool, a reusable procedure, or the model’s weights. Criticizing an answer and trying again can help finish today’s task. Persistent improvement changes what happens when tomorrow’s task begins.

OpenAI’s Tax AI deployment turns practitioner corrections into evaluated product changes, making the quality of this feedback loop a concrete engineering problem.

The interesting question is how to turn those failures into useful practice, then retain what helps on future tasks.

A retry loops inside one task, while a checked version change is loaded by a later task.
Persistence requires a retained artifact and a later execution that uses it.

The answer has two paths. One repairs the agent system: its code, instructions, tools, retrieval, or stored skills. The other improves the model using selected demonstrations or rewards from fresh attempts. Both need a way to decide whether a proposed change actually helps.

That is where environments become important. An environment supplies a task, a starting state, tools that change that state, and a way to evaluate the consequences. It can produce new experience when the agent tries a different action. A saved conversation contains only the experience that already happened.

This distinction changes what we should scale. More failure logs can reveal recurring defects. More executable environments can let an agent practice different decisions. Neither produces improvement until an update is made and tested on future behavior.

A useful first check is to ask what survives the run. A reflection left in a discarded conversation changes nothing for the next user. A saved skill changes future behavior only if another run retrieves and executes it. A newly trained checkpoint changes nothing until the application serves it. Persistence needs an actual connection to the next execution.

We will follow that connection through traces, three ways to construct executable environments, training, curriculum, and deployment.

Contents

  • What persists after a run

  • Turn traces into testable failures

  • An environment produces new experience

  • How executable environments are built

  • Check what the environment rewards

  • Convert experience into an agent update

  • Build a curriculum from agent outcomes

  • Promote improvements on independent evidence

User's avatar

Continue reading this post for free, courtesy of Damien Benveniste.

Or purchase a paid subscription.
© 2026 AiEdge · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture