Research note · 2026-10-06
Most AI failures in production are plain ones
The failures that hurt most in deployed systems are rarely exotic. An instruction hidden in a web page changes what an agent does. A tool has more permission than the task needed. A long chain of steps drifts away from the original request.
These are engineering problems and they can be studied as such. Each can be written as a reproducible case, with the boundary that was crossed and the step where it happened.
The first thing we want to build is a map of what an agent is allowed to touch for a given task, and a way to test that map under pressure rather than only under normal use.
Publishing failure details can help attackers as well as defenders. Each case needs a review of what is safe to release and what should be described only in general terms.
The goal is modest. A catalogue of failures and defences that another team can use in a week is worth more than a framework nobody applies.
Related project: Bounds