In brief
Capability is concentrating wherever compute is. Lean asks how much a smaller, cheaper or local system can do when it is designed around the task.
Illustration of the idea. Not a result.
The problem
Most of the world builds with limited compute. Research that only works at the frontier leaves them out.
Capability is concentrating wherever compute is. That leaves most of the world with systems that are too costly, too distant or too slow to depend on.
Approach
How we are going about it.
- 01
Fit the model to the task
Measure what a task needs instead of defaulting to the largest model available.
- 02
Run where the data is
Study local and on-device operation so capability is not reserved for whoever can pay for the largest cluster.
- 03
Report the cost honestly
Every result is stated with what it costs to run, not only how well it scores.
Open questions
- Which tasks genuinely need a large model, and which only seem to?
- How do you share work between a small local model and a larger remote one without leaking data?
- What is a fair way to account for the cost of a retry?
What we aim to publish
- Cost-attribution methods for agent workloads
- Comparisons of small and large systems on field tasks
- Reference designs for local-first assistants
Limits. Smaller systems have real ceilings. Lean will say where they are.
Landscape
What already exists, and where we start.
Scaling-law work explains why compute concentrates. Distillation, adaptation, quantisation and local runtimes show what can be made smaller. We want to ask the task-first question: what is the cheapest system that is good enough for this job, and how would you know?
- PaperKaplan et al. (2020). Scaling Laws for Neural Language Models.
Showed loss falling predictably with model size, data and compute. The starting point of the compute-first view.
- PaperHoffmann et al. (2022). Training Compute-Optimal Large Language Models.
Revised how model size and training data should be balanced for a fixed compute budget.
- PaperHinton, Vinyals and Dean (2015). Distilling the Knowledge in a Neural Network.
Trains a small model to imitate a large one.
- PaperHu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models.
Adapts a large model by training a small number of added parameters.
- PaperDettmers et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs.
Fine-tunes a quantised model on modest hardware.
- PaperFrantar et al. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.
Compresses model weights to a few bits with limited loss in quality.
- Toolllama.cpp.
An open-source runtime for running language models on ordinary local hardware.
Notes