A self-improving agent uses evidence from earlier tasks to change something retained for later tasks, then demonstrates better behavior on new runs. The change might be a prompt, a tool, a reusable procedure, or the model’s weights. Criticizing an answer and trying again can help finish today’s task. Persistent improvement changes what happens when tomorrow’s task begins.
OpenAI’s Tax AI deployment turns practitioner corrections into evaluated product changes, making the quality of this feedback loop a concrete engineering problem.
The interesting question is how to turn those failures into useful practice, then retain what helps on future tasks.
The answer has two paths. One repairs the agent system: its code, instructions, tools, retrieval, or stored skills. The other improves the model using selected demonstrations or rewards from fresh attempts. Both need a way to decide whether a proposed change actually helps.
That is where environments become important. An environment supplies a task, a starting state, tools that change that state, and a way to evaluate the consequences. It can produce new experience when the agent tries a different action. A saved conversation contains only the experience that already happened.
This distinction changes what we should scale. More failure logs can reveal recurring defects. More executable environments can let an agent practice different decisions. Neither produces improvement until an update is made and tested on future behavior.
A useful first check is to ask what survives the run. A reflection left in a discarded conversation changes nothing for the next user. A saved skill changes future behavior only if another run retrieves and executes it. A newly trained checkpoint changes nothing until the application serves it. Persistence needs an actual connection to the next execution.
We will follow that connection through traces, three ways to construct executable environments, training, curriculum, and deployment.
Contents
What persists after a run
Turn traces into testable failures
An environment produces new experience
How executable environments are built
Check what the environment rewards
Convert experience into an agent update
Build a curriculum from agent outcomes
Promote improvements on independent evidence




