In brief
Most real failures are plain: an injected instruction, a tool with too much permission, a long chain that drifts. Bounds studies exactly where systems break and publishes what is safe to share.
Illustration of the idea. Not a result.
The problem
Safety research that stays abstract does not help the team shipping next quarter. Bounds is applied and modest on purpose.
The failures that matter most in deployed systems are unglamorous: an injected instruction in a document, a permission that was broader than intended, a chained workflow that did something nobody asked for.
Approach
How we are going about it.
- 01
Map the boundary
Write down what an agent is allowed to touch, and test that boundary under pressure, not just under normal use.
- 02
Study real failure modes
Prompt injection, permission creep and workflow drift are studied as engineering problems with reproducible cases.
- 03
Publish what saves time
The aim is a catalogue of failures and defences that another team can use in a week.
Open questions
- How do we publish useful defensive detail without handing over working attacks?
- How do protections behave when workflows are long and tools are chained?
- Which findings are about one system, and which generalise?
What we aim to publish
- A public taxonomy of failure modes
- Replicated findings across more than one model provider
- Guidance for builders on scoped permissions
Limits. Publishing failure details can help attackers. Each item is reviewed for what it is safe to release.
Landscape
What already exists, and where we start.
The OWASP list and NIST framework say what to worry about. Greshake et al. showed one concrete failure class. We want reproducible cases at the level of a single permission or a single step, so another team can test their own system the same way.
- PaperGreshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.
Described indirect prompt injection against applications that feed untrusted content to a model.
- StandardOWASP Top 10 for Large Language Model Applications.
A community list of the most common security risks in LLM applications.
- StandardNIST (2023). AI Risk Management Framework (AI RMF 1.0).
A voluntary US framework for mapping, measuring and managing AI risk.
- StandardISO/IEC 42001:2023. Information technology, Artificial intelligence, Management system.
A management-system standard for organisations that develop or use AI.
Notes