Pipeline
From a question to something you can use.
Public work, work in development, concepts we are scoping and ideas for later. Codenames are our own working names. Stages show real maturity: nothing here is a result.
Public now
Open to read today.
Knowledge baseLibraryA growing reading list of the public work our projects build on.
WritingResearch notesShort notes and open questions on the six projects.
In development
Active early research. No results yet.
Flagship projectReplayRecorded, replayable agent runs.
Flagship projectRecallMemory with a source, an owner and an expiry.
Flagship projectProofEvaluations a field expert would sign off on.
Flagship projectBoundsReproducible cases of plain AI failures.
Flagship projectLeanCapable systems on smaller compute.
MethodFieldworkStarting each program from a practitioner's problem.
Scoping next
Defined as a concept, not started.
One shared record format across replay, memory and evaluation, so a run, a memory and a test result can point at each other.
A way to report disagreement between several experts on the same output instead of averaging it away.
A local-first memory store that keeps provenance on the device it was created on.
A sandbox for testing what an agent can touch, one permission at a time.
Internal testing
Appears here only when something is genuinely being tested.
Nothing at this stage yet.
Preparing for release
Appears here only when something is genuinely close.
Nothing at this stage yet.
Long-term ideas
Speculative. May change or never happen.
Small task-specific systems derived from recorded runs. We do not know if this works.
An open evaluation commons contributed to by practitioners across fields.
A knowledge base linking what each field program learns to the others.