In brief
An agent run is usually a stream of text that disappears. Replay treats a run as a recorded object, so a failure can be re-run, changed at one step and compared.
Illustration of the idea. Not a result.
The problem
When an agent does something wrong in a long workflow, the usual evidence is a transcript. A transcript shows what was said, not what the system actually did or what would have happened if one input had been different.
Most agent systems today are judged by their final answer. When the answer is wrong, the run that produced it is gone: nothing to replay, nothing to compare, nothing to learn from.
Approach
How we are going about it.
- 01
Record at the edge
Everything non-deterministic that enters a run (model output, tool results, clock, randomness) is written down as it happens. The rest of the run is then a pure function of those records.
- 02
Replay as an operation
A recorded run can be re-executed against its own records. Change one record and re-execute, and the two runs can be compared step by step.
- 03
Small units of blame
A cause should point at one step, not at a whole conversation. Replay keeps the unit of debugging small enough to reason about.
Open questions
- How much of a real run can be made deterministic without hiding the behaviour you care about?
- What is the cheapest record that still lets a run be replayed faithfully?
- How should a replay be presented to someone who did not write the agent?
What we aim to publish
- A public specification for recorded runs
- A reference replay tool
- Worked examples of diagnosing a failure by diffing two runs
Limits. Some behaviour cannot be recorded faithfully, such as a tool with hidden state. Replay will mark those steps as unreplayable instead of pretending.
Landscape
What already exists, and where we start.
rr shows that record and replay works for ordinary programs. Agent runs add a model that is not deterministic and tools with outside effects, so the thing to record is different. We do not yet know how much of a real run can be replayed faithfully. That is the research question, not a solved part.
- Toolrr: record and replay debugger.
Records a program execution and replays it deterministically. The closest mature idea to run-level replay for agents.
- PaperGreshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.
Described indirect prompt injection against applications that feed untrusted content to a model.
- StandardNIST (2023). AI Risk Management Framework (AI RMF 1.0).
A voluntary US framework for mapping, measuring and managing AI risk.
- RegulationEuropean Union (2024). Regulation (EU) 2024/1689, the Artificial Intelligence Act.
The EU's risk-based law for AI systems. Sets record-keeping and oversight duties for high-risk uses.
Notes