Glossary
Terms, in plain words.
- Agent
- A software system that takes a goal, decides on steps, and uses tools to carry them out.
- Agent run
- One complete attempt by an agent at a goal, from the first input to the final output.
- Attribution (cost)
- Assigning spend to the step or feature that caused it.
- Audit trail
- A tamper-evident record of what a system accessed and did.
- Benchmark saturation
- When top systems score so highly on a test that it no longer distinguishes between them.
- Context
- The information an assistant has available while working on a task.
- Determinism
- The property that the same inputs produce the same output every time.
- Evaluation
- A structured way of measuring how well a system does a defined piece of work.
- Explainability
- The ability to show why a system produced a particular result.
- Fine-tuning
- Further training a model on a narrower set of examples.
- Frontier (of a run)
- The set of recorded values entering a run from outside it, such as model outputs and tool results.
- Guardrail
- A rule or check that limits what a system can do.
- Hallucination
- Output that sounds plausible but is not supported by the source material.
- Local-first
- Designed so the core work happens on the user’s own device.
- Prompt injection
- Instructions hidden in content an assistant reads, intended to change what it does.
- Provenance
- The recorded origin and history of a piece of data or content.
- Red teaming
- Deliberately trying to make a system fail, in order to find weaknesses.
- Replay
- Re-executing a recorded run against its own stored records to reproduce it.
- Retrieval
- Looking up relevant material at the time a question is asked.
- Scoped access
- Permission limited to what a particular task needs, for as long as it needs it.