Library
The work we build on.
A growing reading list. Every entry is a public work. The notes are our own reading, not the authors' claims.
Make every agent run something you can pause, inspect and replay.
- Toolrr: record and replay debugger.
Records a program execution and replays it deterministically. The closest mature idea to run-level replay for agents.
- PaperGreshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.
Described indirect prompt injection against applications that feed untrusted content to a model.
- StandardNIST (2023). AI Risk Management Framework (AI RMF 1.0).
A voluntary US framework for mapping, measuring and managing AI risk.
- RegulationEuropean Union (2024). Regulation (EU) 2024/1689, the Artificial Intelligence Act.
The EU's risk-based law for AI systems. Sets record-keeping and oversight duties for high-risk uses.
Memory an assistant can prove, correct and forget.
- PaperPacker et al. (2023). MemGPT: Towards LLMs as Operating Systems.
Manages a model's context like an operating system pages memory. Treats memory as a systems problem.
- PaperLewis et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.
Combined a generator with a retrieved document store. The base of most retrieval-based memory today.
- PaperGebru et al. (2018). Datasheets for Datasets.
Proposes documenting a dataset's origin, composition and intended use. A direct ancestor of source-bearing memory.
- StandardW3C (2013). PROV Overview.
A W3C family of specifications for representing provenance: who or what produced a piece of data, and how.
Tests a field expert would actually sign off on.
- PaperHendrycks et al. (2020). Measuring Massive Multitask Language Understanding.
A widely used multiple-choice benchmark across 57 subjects. A useful case study in benchmark saturation.
- PaperLiang et al. (2022). Holistic Evaluation of Language Models.
A broad evaluation framework from Stanford CRFM that reports many metrics across many scenarios.
- PaperJimenez et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Evaluates models on real software issues instead of synthetic questions. Close to our idea of starting from real work.
- PaperMitchell et al. (2019). Model Cards for Model Reporting.
A documentation format for what a model is for and where it fails.
The unglamorous failures that break deployed AI.
- PaperGreshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.
Described indirect prompt injection against applications that feed untrusted content to a model.
- StandardOWASP Top 10 for Large Language Model Applications.
A community list of the most common security risks in LLM applications.
- StandardNIST (2023). AI Risk Management Framework (AI RMF 1.0).
A voluntary US framework for mapping, measuring and managing AI risk.
- StandardISO/IEC 42001:2023. Information technology, Artificial intelligence, Management system.
A management-system standard for organisations that develop or use AI.
Capable systems that do not need the biggest cluster.
- PaperKaplan et al. (2020). Scaling Laws for Neural Language Models.
Showed loss falling predictably with model size, data and compute. The starting point of the compute-first view.
- PaperHoffmann et al. (2022). Training Compute-Optimal Large Language Models.
Revised how model size and training data should be balanced for a fixed compute budget.
- PaperHinton, Vinyals and Dean (2015). Distilling the Knowledge in a Neural Network.
Trains a small model to imitate a large one.
- PaperHu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models.
Adapts a large model by training a small number of added parameters.
- PaperDettmers et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs.
Fine-tunes a quantised model on modest hardware.
- PaperFrantar et al. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.
Compresses model weights to a few bits with limited loss in quality.
- Toolllama.cpp.
An open-source runtime for running language models on ordinary local hardware.
Start every program from a problem a practitioner can state.
- StandardNIST (2023). AI Risk Management Framework (AI RMF 1.0).
A voluntary US framework for mapping, measuring and managing AI risk.
- RegulationEuropean Union (2024). Regulation (EU) 2024/1689, the Artificial Intelligence Act.
The EU's risk-based law for AI systems. Sets record-keeping and oversight duties for high-risk uses.
- StandardISO/IEC 42001:2023. Information technology, Artificial intelligence, Management system.
A management-system standard for organisations that develop or use AI.
- PaperMitchell et al. (2019). Model Cards for Model Reporting.
A documentation format for what a model is for and where it fails.