Skip to main content

AI systems lab / Clarence memory / 2026

How Clarence remembers

Clarence is the AI system I work with every day. Since March 2026 I have been building the part that decides what it remembers. This page shows how that memory works, how I measured it, what broke along the way, and what I am fixing next.

I designed the memory system, set its rules, and approved every structural change. Clarence, Hermes (the agent software Clarence runs on) and Claude Code wrote most of the code under that direction. Figures come from the live database on September 29, 2026.

6,173
memory records
15,623
facts about people, projects and tools
60,476
indexed passages from my documents
1
agent allowed to write memory
19→34
first-place answers of 40, after reranking
7 mo
of building, 9 rebuilds, 17+ audits

As of September 29, 2026. About half of all records are archived; nothing is deleted.

Section / 01

Why I care about memory

Every AI tool I used before Clarence started from zero. I explained my projects, my constraints and my past decisions again in each new window. In March 2026 I began building a system that would keep that context between sessions.

The goal grew. I want a knowledge system that learns with me as I read, work and build. Memory became the hardest part of the whole project. Storing things was easy. The hard questions were which records deserve to be kept, which version is current, who may change it, and whether the right record comes back when an agent needs it.

"Total recall … I want to be able to trust that this system knows me."My design brief for Reed, August 26, 2026

Section / 02

Six places knowledge lives

Clarence has no single memory. Each layer answers a different question, and much of the work has been deciding which layer holds the truth for which question.

Memory layerscurrent

Always-loaded notessmall, in every promptA few thousand characters each agent carries into every conversation. Fast, tiny, and full.
Session historyper agentEvery conversation, searchable. Answers "what happened, in order."
Clarence databaseSQLite + vectorsMemory records, facts about entities, a relationship graph, my notes and documents split into passages. Every active item is indexed by meaning. Only Reed writes here.
System wiki232 notesCompiled, human-readable truth about each project. Answers "what is true and how does it connect." Reed keeps its activity logs and timelines.
My notesabout 2,000 notesMy own writing. The database indexes it. Only I promote a note into my permanent collection.
Specialist notesper toolClaude Code keeps its own notes about how I work. They are separate from the database.
The public knowledge graph shows the shape of the database without its contents.

Section / 03

One writer

From March to August several agents could write to memory. They saved conflicting versions of the same facts, and nobody could say which one was current. An audit on August 23 found 85% of active memories dated from March and 97% past their review date. Fields meant to track currency existed, and nothing used them.

On August 26 I designed a different arrangement. Reed is an agent whose only job is keeping memory. Every other agent, Clarence included, submits what it learned to a queue. Reed checks each submission, saves it with a stamp naming who sent it, from which machine and session, and when, and logs the decision.

Nothing is deleted. An outdated record is marked as replaced and points to the record that replaced it. A change to a fact about a person waits for my approval.

Write pathlive since September 2026

Agents

Clarence, specialist agents and Claude sessions on three devices finish work or learn something.

Queue

Each submission is a file with its origin stamp. Any agent can submit; none can write.

Reed

Checks, saves, applies corrections to facts, logs every decision. Runs every ten minutes on a local GPU.

Readers

Every agent searches the same database. A daily test record proves the full loop still works.

1,010 submissions landed and 3 were rejected, each with a written reason, through September 28.

Reed started on August 31, on a GPU I installed that day. On September 8 a seven-day check confirmed that every write in the window came from Reed. Reed's judgment, nightly reflection and weekly review all run on local hardware, so memory upkeep spends no API budget.

Section / 04

Finding the right memory

A stored fact is useless if the agent never finds it. Clarence searches by meaning: a small open-source model turns every memory and fact into a list of 384 numbers (an embedding), and search returns the nearest matches. A second model, called a reranker, then reads the question beside each candidate and reorders them.

I tested each change against a fixed set of questions with known answers.

What the tests showedshare of questions answered

Reranking, first-place answersApr 2026 · 19 → 34 of 40
Routing questions to the wiki, top 20May 2026 · 23 → 73 of 75
Local reranker on my own GPU, top resultsSep 2026 · 29 → 37 of 40
Keyword ranking blended with meaning, top 30Sep 2026 · 23 → 28 of 30
beforeafter
Each bar is drawn to its own test's scale. The tests used different question sets and are not comparable with one another.

Two results changed how I build. In April, a larger embedding model lost to the small one already in use, 14 questions to 17, so I kept the small one. In May, clearing a backlog of unprocessed documents moved the score by zero, while sending questions to the wiki first took it from 23 to 73 of 75. Where a question goes mattered more than how much data sat behind it.

Section / 05

Seven months, nine rebuilds

  1. March 2026Markdown files and a first database

    Daily notes, a curated memory file and a small SQLite database. Three databases were merged into one on March 25 with 171 memories. Bulk imports took it past 2,300 memories and 9,000 facts within a week.

  2. April 2026Cleanup and first measurements

    2,363 low-value items archived. A bug found where search compared vectors of two different sizes and returned nothing, with no error. Clarence moved to Hermes. The system wiki began as the home for compiled truth.

  3. May 2026Routing beats volume

    The 75-question test. A read-only memory connector so other AI tools could consult Clarence's memory without changing it.

  4. June to July 2026Rebuild and audit

    A profile retirement silently orphaned every memory maintenance job. A July 3 audit found it and reclaimed about 15 GB, archived 1,971 bulk records and merged 37 duplicate entities.

  5. August 2026Cost, then a custodian

    A token-cost investigation, a daily memory test, and three audits that agreed the plumbing was healthy and the content was stale. Reed was designed.

  6. September 2026One writer, local models, an automated wiki

    Reed became the only writer on September 1. A local reranker went live. Reed began writing the wiki's activity logs, timelines and a daily record.

Section / 06

What broke

Most failures had the same shape: something stopped working and nothing said so.

WhenWhat happenedScale
Mar 29Thirteen scheduled jobs on one model invented their tool results and reported them as done.every run
to Apr 1Semantic search returned nothing because the query and the index used different vector sizes.silent
Jun 16An oversized prompt rode on every call.51.9M tokens in a day
Jun 25Retiring an old profile orphaned every memory maintenance job.9 days unnoticed
Aug 3 to 17Long sessions re-read their own cached history on every call.1.56B tokens in 2 weeks; about $210 to $240 a week at API rates, estimated
Aug 22A full working day with 553 tool calls and zero memory lookups.1 day
Aug 26A failed build that consumed 70 million tokens left no record in memory.70M tokens
to Sep 25Notes sent from other tools went to an address that no longer existed and were lost.all of them
Sep 25 to 28The hook that loads memory at the start of a conversation stopped running on most sessions.3 days

Rules I build by now

  • One writer for memory.Came from months of agents overwriting each other.
  • Archive; never delete.Came from records disappearing with no trace of why.
  • Every scheduled job has an owner and a health check.Came from the June orphaned jobs.
  • Prove it works every morning.A test record is written and searched for daily.
  • Capture automatically.Came from the 70-million-token build that left nothing behind.
  • Keep the negative results.The tests that showed no gain changed more decisions than the ones that did.

Section / 07

Why search alone is not enough

In May I built an interactive case to show the gap between finding a memory and being allowed to act on it. The example is small. A request to update and publish a portfolio page runs into an old approval and a newer standing rule. It is kept here, updated for the current system.

A normal request hides authority to act.

The task is simple on the surface: update the Clarence case study and publish it. The word publish changes the memory bar because a public action needs permission, not just context.

What the agent believesI need deployment context before I touch production.
What the memory rules sayProposed gate: public action detected. Only a standing rule should be allowed to govern this action.

Pending. Retrieve evidence, then check whether it has authority to act.

proposed: risk.classify action=public_deploy gate=memory_authority

old_deploy_approval
says
"Looks good. Go ahead and deploy this one."
status
superseded built
kind
event built
replaced by
preview_required built
scope
one earlier deploy proposed
may govern
nothing proposed
preview_required
says
"For portfolio changes, show me a preview before deploying."
status
active built
kind
standing rule built
expires
never built
may govern
portfolio deploys proposed
requires
preview, wait, then deploy proposed

Fields marked built exist on every record in the database today: status, kind, replacement pointer, confidence, review date, expiry date and author. Fields marked proposed describe the next layer: what a record is allowed to decide, and what action it requires. Search can rank a record as relevant. It cannot say whether that record has the authority to decide an action.

Section / 08

Where it stood on September 29, 2026

The machinery works. The content has not caught up. Every week Reed asks the system ten questions about me and scores the answers. It scored 9 of 10 on September 13 and 5 of 10 on September 27. In each miss I have traced, the right record was stored and lost the ranking.

WorkingOne writer, full provenance, no deletions, daily memory test, nightly backup, every active item indexed.
Too openReed checks that a submission is well formed and saves it. 283 of 634 recent session summaries came from automated runs with no real conversation, and they crowd out better records.
IncompleteA correction to a memory record describes what it replaces but leaves the old record active, so search can still return it first.
UnevenThe best search blend lives only in the start-of-conversation hook, which runs for two of eight agents and failed silently last week.
UnreviewedReed logs every decision for me to review. I have reviewed 6 of 3,335.
Thin graph86% of entities have no recorded relationship to anything else.

Section / 09

Next steps identified on September 29, 2026

The next month is curation. None of it needs new infrastructure.

  1. A quality check at Reed's door that skips empty automation summaries, catches near-duplicates and matches names to known entities before saving.
  2. Corrections that mark the old memory record as replaced.
  3. The keyword-plus-meaning search blend moved into the shared search server so every agent gets it.
  4. The start-of-conversation hook on every agent, with a morning alert when it goes quiet.
  5. Always-loaded notes rebuilt by Reed from the database, ranked by importance, so they never overflow.
  6. A fixed question set I write myself, so the weekly score measures memory against my answers.
  7. A ten-minute weekly review, so my decisions become the record that trains Reed's judgment.

Further out, I want to package the memory server and the submission queue as open-source tools, and to make the whole system usable from a phone. The goal is the one I started with: a system that learns with me and that other people could use to do the same.

Earlier pagesThis page draws on two earlier pages. Memory Governance (May 6, 2026) introduced the interactive case in section 07. How Clarence Remembers (August 25, 2026) was a static architecture snapshot; its component inventory and source hierarchy are folded into sections 02 and 03. Both old addresses remain available; redirects and site-wide link changes are a separate follow-up.