Mynd LabsResearchAll of Mynd Labs ↗
Research division
CollaboratePressContact

Research note · 2026-10-06

An agent run should be an object, not a transcript

A transcript tells you what an agent said. It does not tell you what the system did, in what order, with which inputs, or what would have happened if one input had been different. When an agent does something wrong inside a long workflow, a transcript is a weak piece of evidence.

The alternative is to treat a run as a recorded object. Every source of non-determinism at the edge of the run is written down: the model output, the tool results, the time, the random seed. Once those are recorded, the rest of the run is a pure function of them.

That one change gives you three operations that are hard today. You can replay a run exactly. You can change one record and replay, to ask what would have happened. And you can compare two runs step by step and find the first place they diverge.

There are real limits. A tool with hidden state cannot be replayed faithfully, and a very long run produces a lot of records. We think the honest answer is to mark the steps that cannot be replayed and to keep the unit of debugging small, so a cause points at one step instead of a whole conversation.

This note describes a direction, not a result. The open question we care about most is the cheapest record that still lets a run be replayed faithfully.

Related project: Replay