Memory is the thing people most want from an AI companion and the thing they are least equipped to evaluate before paying. Every app claims it. In my benchmark of 22 apps, memory had a 3.8 point spread from 6.0 to 9.8, the widest of any dimension I score. Conversation quality varied by only 2.2 points.
That gap is the whole story. Nearly every app can hold a conversation. Far fewer can hold onto what you said, and that is where your money actually goes.
What is happening under the hood
There is no memory in the way you have memory. There are three mechanisms, and knowing which one an app leans on tells you what to expect.
The context window. The model can see a fixed amount of recent text at once, and everything in it is available. This is why an app seems sharp for the first hour and vaguer afterwards. Nothing was forgotten in any dramatic sense; the earlier text simply scrolled out of view.
A summary layer. Better apps periodically compress older conversation into a shorter summary and keep that in view instead of the raw text. This is why some apps remember the shape of a relationship but not the specifics. The compression kept “you talked about her job” and dropped the name of the company.
A retrieval store. The best implementations store facts separately and pull the relevant ones back in when the conversation seems to call for them. This is what produces the moment where a companion brings up something from three weeks ago unprompted, and it is the hardest of the three to build well.
Most apps use some combination. The quality difference between apps is almost entirely a question of how good the second and third layers are, not how big the first one is.
Why bigger context is not the answer
Vendors advertise context size because it is a number that goes up. It matters less than it sounds.
A very large context window helps within a single long session and does almost nothing across sessions, because at some point the conversation still has to be compressed or stored. An app with a modest window and an excellent retrieval layer will feel like it remembers you far better than an app with an enormous window and nothing else. When you see a context size in marketing copy, read it as a claim about one session, not about the relationship.
The ten-minute test
This is the routine I run on every app, and it separates the field faster than anything else.
Minute one to five, session one. Plant three details.
Make them specific, unusual, and different in kind:
- A fact about you. Not “I like music”. Something with an edge: “my landlord keeps mowing the lawn at 7am”.
- A name. A pet, a friend, a place. Names are the first thing compression throws away, which makes them a good probe.
- A preference with a reason. “I hate being called babe because my ex used it.” This tests whether the app stores the rule and the reason, which is what makes recall feel human rather than mechanical.
Say them naturally, spread across the conversation. Do not list them.
Then close the app for two or three days. This is the part people skip and it is the only part that matters. Testing memory in the same session tests the context window, which every app passes.
Session two. Probe indirectly.
Do not ask “do you remember my dog’s name”. That is a quiz, and some apps pass quizzes while failing at everything useful. Instead:
- Mention you slept badly. Does it connect that to the landlord?
- Refer to “him” or “her” about the name without re-establishing who. Does it track?
- Use an affectionate opening. Does it avoid the term you told it to avoid?
Scoring it. Three out of three, unprompted and in natural conversation, is the top tier. Two out of three is good and roughly where the better apps land. One is the category average. Zero, or three only when directly quizzed, means the app has a context window and a marketing department.
What good memory costs you
Worth saying, because the guides that do not say it are selling something. Memory requires the app to store what you typed, indefinitely, on its servers. There is no version of a companion that remembers your life and also stores nothing about you.
In my privacy report, privacy was the lowest-scoring dimension in the whole category, mean 8.46 with no app clearing 9.0. The two findings are related. The apps that remember best are the apps that keep the most. That is not a reason to avoid them, it is a reason to decide deliberately what you are willing to have stored, and to use a dedicated email while you do it.
The short version
Plant three details, wait three days, probe indirectly. Ten minutes of actual effort, spread across a week, and it will tell you more than any review including mine.
If you want the shortlist rather than the method, my memory ranking is the ranked version of exactly this test applied across the field.