← Back to the archive
Essay/Full text

An Archive Is Not a Memory

Why using an AI agent’s memory should change it — and what separates retrieval from learning.

Think of someone you know well, and notice what comes back. A room. A feeling in your body. An argument with someone else entirely. The memories arrive sideways, by association, often sharing not a single word with the thought that triggered them.

Now look at how we build “memory” for AI agents: chop the past into chunks of text, index them, and when a question comes in, fetch the chunks that look most similar. It works well enough that we’ve all agreed to call it memory. But it’s an archive with a good search engine, and an archive answers a different question than a memory does. An archive asks: what looks like the query? A memory asks: what is connected to this, through everything that came before?

There’s a second difference, and it’s the one this essay is about. An archive is unchanged by being searched. A memory isn’t. Recall is a write operation, not a read - every act of remembering makes some paths easier to find next time and lets others fade.

The unit is wrong

Here’s a typical stored chunk:

Danny was late to the meeting. He apologized. We postponed the launch. I left the call frustrated.

One chunk, one entry in the index - but four separate claims, each belonging to a different web. The lateness bears on whether Danny is reliable. The postponement, on the project timeline. The frustration belongs to a pattern I have around professional uncertainty. And the apology is evidence against a belief I formed later - that Danny doesn’t respect me.

Fused into one blob, none of these can be reached alone - you can’t retrieve the apology without dragging the frustration along. So the unit should be the single claim, tied to the moment that produced it. But a pile of claims is still a pile. The meaning was never inside the sentences; it lives between them. This caused that. This contradicts that. This is the kind of thing that keeps happening.

Two numbers, not one

In your head, connections are not equal: some are worn smooth by use, others exist and are never walked again. So here’s the claim I care about: every link in a memory should carry two numbers, not one. The first is confidence - how likely the connection is to be true. The second is path value - how useful the link is to walk when answering a certain kind of question.

They diverge constantly. “Danny works at the company” is about as certain as a stored fact gets, and it helps almost no question you’ll ever ask. “This feels like that episode with a different manager” asserts no fact at all - and sometimes it’s exactly the route to the insight. Truth and usefulness are different axes.

So: similarity proposes a connection. Evidence justifies it. Only outcomes -walking it actually helped - decide whether it’s worth following again. That’s how the desire paths form: trails appearing wherever walking actually happened, not wherever someone planned a road.

What walking buys you

Say a user asks why they’re anxious about the meeting with Danny. No stored sentence says that. Similarity search finds the obvious seeds - Danny, meetings, anxiety - and from there the system walks: Danny → criticism during the previous project → fear of being unprepared → recurring stress before evaluations → a similar episode with a different manager.

A path like that surfaces a pattern no paragraph states - the thing an old friend knows about you without ever having been told. If surfacing it genuinely helped, the links along that route should get slightly easier to reach next time. That’s the whole proposal: retrieval that leaves footprints.

The honest part

The trap hides inside “genuinely helped.” An answer succeeds, and behind it sat twenty claims connected by thirty links - which link earned it? Settle that by asking the model that wrote the answer, and you’ve built a memory that grades its own homework. The honest version is a question you can ask yourself too: what would I have concluded if I’d never made this association? Remove one branch of the memories behind an answer and see whether it still stands - judged by the next decision, not by the system’s opinion of itself. Because a memory that strengthens whatever it already likes isn’t learning. It’s ruminating.

It fails the way we fail

A valuable path gets walked more, which makes it stronger, which gets it walked more - the machine version of telling yourself the same story about someone for years. The cure is the same boring one it is for people: deliberately walking unfamiliar routes, re-checking old conclusions.

Which leads to the corollary nobody likes: in a system like this, forgetting is a feature you have to build, not a failure you tolerate. A path that stops earning its keep should fade - decay is how the map stays honest about which routes still lead anywhere. A memory that can’t forget doesn’t remember better. It just recites.

And one failure isn’t technical at all. A system that stores inferred claims about a person’s fears is building a psychological dossier, and “it’s just memory” doesn’t cover that. The person should be able to see - and contest - what the system believes about them. An inference you never made about yourself, surfaced at the wrong moment, can hurt even when it’s accurate. Especially then.

Memory as the history of its own use

The standard question is what to store and how to find it later. The question pulling at me changes one word: how should using a memory change the odds of reaching it again? Get that right and the system starts carrying a record of its own reasoning - which paths answered real questions, which conclusion quietly expired in March. The records stay fixed. The routes through them keep moving.

Maybe none of this beats a very long context window and a good filing cabinet - that’s an empirical question, and the full mechanism and an experiment for settling it are here:

The full mechanism and proposed experiment

But the point underneath isn’t empirical. We haven’t built memory for our agents. We’ve built retrieval: read-only memory. And the defining property of the real thing is that reading it writes it.