Mynd LabsResearchAll of Mynd Labs ↗
Research division
CollaboratePressContact

Library

The work we build on.

A growing reading list. Every entry is a public work. The notes are our own reading, not the authors' claims.

Replay

Make every agent run something you can pause, inspect and replay.

  1. Tool
    rr: record and replay debugger.

    Records a program execution and replays it deterministically. The closest mature idea to run-level replay for agents.

  2. Paper
    Greshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.

    Described indirect prompt injection against applications that feed untrusted content to a model.

  3. Standard
    NIST (2023). AI Risk Management Framework (AI RMF 1.0).

    A voluntary US framework for mapping, measuring and managing AI risk.

  4. Regulation
    European Union (2024). Regulation (EU) 2024/1689, the Artificial Intelligence Act.

    The EU's risk-based law for AI systems. Sets record-keeping and oversight duties for high-risk uses.

Recall

Memory an assistant can prove, correct and forget.

  1. Paper
    Packer et al. (2023). MemGPT: Towards LLMs as Operating Systems.

    Manages a model's context like an operating system pages memory. Treats memory as a systems problem.

  2. Paper
    Lewis et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

    Combined a generator with a retrieved document store. The base of most retrieval-based memory today.

  3. Paper
    Gebru et al. (2018). Datasheets for Datasets.

    Proposes documenting a dataset's origin, composition and intended use. A direct ancestor of source-bearing memory.

  4. Standard
    W3C (2013). PROV Overview.

    A W3C family of specifications for representing provenance: who or what produced a piece of data, and how.

Proof

Tests a field expert would actually sign off on.

  1. Paper
    Hendrycks et al. (2020). Measuring Massive Multitask Language Understanding.

    A widely used multiple-choice benchmark across 57 subjects. A useful case study in benchmark saturation.

  2. Paper
    Liang et al. (2022). Holistic Evaluation of Language Models.

    A broad evaluation framework from Stanford CRFM that reports many metrics across many scenarios.

  3. Paper
    Jimenez et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    Evaluates models on real software issues instead of synthetic questions. Close to our idea of starting from real work.

  4. Paper
    Mitchell et al. (2019). Model Cards for Model Reporting.

    A documentation format for what a model is for and where it fails.

Bounds

The unglamorous failures that break deployed AI.

  1. Paper
    Greshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.

    Described indirect prompt injection against applications that feed untrusted content to a model.

  2. Standard
    OWASP Top 10 for Large Language Model Applications.

    A community list of the most common security risks in LLM applications.

  3. Standard
    NIST (2023). AI Risk Management Framework (AI RMF 1.0).

    A voluntary US framework for mapping, measuring and managing AI risk.

  4. Standard
    ISO/IEC 42001:2023. Information technology, Artificial intelligence, Management system.

    A management-system standard for organisations that develop or use AI.

Lean

Capable systems that do not need the biggest cluster.

  1. Paper
    Kaplan et al. (2020). Scaling Laws for Neural Language Models.

    Showed loss falling predictably with model size, data and compute. The starting point of the compute-first view.

  2. Paper
    Hoffmann et al. (2022). Training Compute-Optimal Large Language Models.

    Revised how model size and training data should be balanced for a fixed compute budget.

  3. Paper
  4. Paper
    Hu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models.

    Adapts a large model by training a small number of added parameters.

  5. Paper
    Dettmers et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs.

    Fine-tunes a quantised model on modest hardware.

  6. Paper
  7. Tool
    llama.cpp.

    An open-source runtime for running language models on ordinary local hardware.

Fieldwork

Start every program from a problem a practitioner can state.

  1. Standard
    NIST (2023). AI Risk Management Framework (AI RMF 1.0).

    A voluntary US framework for mapping, measuring and managing AI risk.

  2. Regulation
    European Union (2024). Regulation (EU) 2024/1689, the Artificial Intelligence Act.

    The EU's risk-based law for AI systems. Sets record-keeping and oversight duties for high-risk uses.

  3. Standard
    ISO/IEC 42001:2023. Information technology, Artificial intelligence, Management system.

    A management-system standard for organisations that develop or use AI.

  4. Paper
    Mitchell et al. (2019). Model Cards for Model Reporting.

    A documentation format for what a model is for and where it fails.