I finished re-scoring my full review corpus this month, 22 apps on the same six dimensions, and one number reframed how I think about the whole category.
Conversation quality across those 22 apps ranges from 7.5 to 9.7. That is a 2.2 point spread. Memory ranges from 6.0 to 9.8, a spread of 3.8 points, and that is only counting the apps where memory exists in a form worth scoring at all. Eight of the 22 did not qualify.
Why that matters
For most of the time I have been covering these apps, conversation quality was the thing worth writing about. Some apps wrote well and some wrote like a support script, and the difference was obvious in the first ten minutes.
That era is over. Every app in my current sample can hold a conversation. The bottom of the field is now competent rather than embarrassing, which is a real achievement and also means the thing everyone optimises their marketing around has stopped being a differentiator.
What has not converged is memory. And memory is the thing that determines whether an app is a companion or a very good chatbot you talk to repeatedly.
The practical consequence
It changes the buying question. “Does it chat well” now has the same answer nearly everywhere, so asking it tells you nothing. “Does it remember” splits the field roughly in half.
It also changes how you should read a free trial. Conversation quality is visible in ten minutes, which is exactly the window a trial is designed to show you. Memory is invisible in ten minutes and only becomes apparent on day three, after the trial has done its job. The dimension that matters most is the one the sales process is structured to hide.
I wrote up the ten-minute test I use in how AI companion memory actually works. The short version: plant three specific details, wait three days, probe indirectly rather than quizzing.
The uncomfortable half of it
Memory requires storage. An app that remembers what you told it in July is an app that still has what you told it in July.
In my privacy report, privacy came out as the weakest dimension in the entire category, mean 8.46 out of 10 with no app clearing 9.0. Those two findings are the same finding viewed from opposite sides. The feature I am telling you to prioritise is the feature that makes the privacy picture worse, and I would rather say that plainly than recommend memory-first apps without mentioning it.
That does not make it a bad trade. It makes it a trade, which is different from how it usually gets sold.
What I am changing
Three things in how I cover this going forward.
Memory gets tested over days rather than within a session, on every review, without exception. A single-session memory check tests the context window and every app passes it.
Reviews say explicitly when an app has no persistent memory worth scoring, rather than quietly omitting the dimension. Eight of 22 is too many to leave implicit.
And the rankings now lead with position rather than a decimal score, because a category where 13 of 22 apps score above 9.0 is one where the second decimal place was pretending to a precision the testing does not support. Where an app sits relative to the others is the honest signal. The main ranking reflects that already.