In brief
An assistant that remembers everything is dangerous, and one that remembers nothing is useless. Recall is about memory with a source, an owner and an expiry.
Illustration of the idea. Not a result.
The problem
People decide whether to trust an assistant by asking how it knows. Most current memory cannot answer that question.
An assistant that remembers everything is dangerous. One that remembers nothing is useless. The real question is what should cross from one task to the next, and who decided.
Approach
How we are going about it.
- 01
Every memory has a source
A remembered fact keeps a pointer to where it came from, so the assistant can show its evidence instead of asserting it.
- 02
Correction is first-class
If a memory is wrong, there is one place to fix it and the fix travels to everything derived from it.
- 03
Forgetting is a feature
Memory can expire, be scoped to a task, or be removed on request, and the system can show that it was removed.
Open questions
- What is the right unit of permission for personal context: a field, a document, a question?
- How do you let a person audit what an assistant used, in language they understand?
- How should memory be shared inside a team without leaking between individuals?
What we aim to publish
- A query model for scoped access to personal context
- A test suite for memory leakage across tasks
- Plain-language audit views for end users
Limits. Provenance does not make a memory true. It makes it checkable. Recall does not try to decide what is true.
Landscape
What already exists, and where we start.
Retrieval and memory-management work focus on getting the right content into a model's context. They say less about where a memory came from, who owns it and when it should go. W3C PROV and datasheets give vocabulary for provenance. Using them for assistant memory is the open step.
- PaperPacker et al. (2023). MemGPT: Towards LLMs as Operating Systems.
Manages a model's context like an operating system pages memory. Treats memory as a systems problem.
- PaperLewis et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.
Combined a generator with a retrieved document store. The base of most retrieval-based memory today.
- PaperGebru et al. (2018). Datasheets for Datasets.
Proposes documenting a dataset's origin, composition and intended use. A direct ancestor of source-bearing memory.
- StandardW3C (2013). PROV Overview.
A W3C family of specifications for representing provenance: who or what produced a piece of data, and how.
Notes