Mynd LabsResearchAll of Mynd Labs ↗
Research division
CollaboratePressContact

Research / Agent runtime

Replay

Make every agent run something you can pause, inspect and replay.

Early researchAgentsReproducibilityTooling

In brief

An agent run is usually a stream of text that disappears. Replay treats a run as a recorded object, so a failure can be re-run, changed at one step and compared.

Recorded runReplay with one record changedfirst divergence

Illustration of the idea. Not a result.

The problem

When an agent does something wrong in a long workflow, the usual evidence is a transcript. A transcript shows what was said, not what the system actually did or what would have happened if one input had been different.

Most agent systems today are judged by their final answer. When the answer is wrong, the run that produced it is gone: nothing to replay, nothing to compare, nothing to learn from.

Approach

How we are going about it.

  1. 01

    Record at the edge

    Everything non-deterministic that enters a run (model output, tool results, clock, randomness) is written down as it happens. The rest of the run is then a pure function of those records.

  2. 02

    Replay as an operation

    A recorded run can be re-executed against its own records. Change one record and re-execute, and the two runs can be compared step by step.

  3. 03

    Small units of blame

    A cause should point at one step, not at a whole conversation. Replay keeps the unit of debugging small enough to reason about.

Open questions

  • How much of a real run can be made deterministic without hiding the behaviour you care about?
  • What is the cheapest record that still lets a run be replayed faithfully?
  • How should a replay be presented to someone who did not write the agent?

What we aim to publish

  • A public specification for recorded runs
  • A reference replay tool
  • Worked examples of diagnosing a failure by diffing two runs

Limits. Some behaviour cannot be recorded faithfully, such as a tool with hidden state. Replay will mark those steps as unreplayable instead of pretending.

Landscape

What already exists, and where we start.

rr shows that record and replay works for ordinary programs. Agent runs add a model that is not deterministic and tools with outside effects, so the thing to record is different. We do not yet know how much of a real run can be replayed faithfully. That is the research question, not a solved part.

  1. Tool
    rr: record and replay debugger.

    Records a program execution and replays it deterministically. The closest mature idea to run-level replay for agents.

  2. Paper
    Greshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.

    Described indirect prompt injection against applications that feed untrusted content to a model.

  3. Standard
    NIST (2023). AI Risk Management Framework (AI RMF 1.0).

    A voluntary US framework for mapping, measuring and managing AI risk.

  4. Regulation
    European Union (2024). Regulation (EU) 2024/1689, the Artificial Intelligence Act.

    The EU's risk-based law for AI systems. Sets record-keeping and oversight duties for high-risk uses.

Work on Replay with us.

hello@myndlabs.tech