2
1 Comment

Every AI Companion App Advertises Memory. I Checked Five and Not One Publishes a Number.

Researchers at Tsinghua University and USTB published a result this year that explains companion app pricing better than any pricing page does. Run a 1 billion parameter model at a 128,000 token context and roughly 90 percent of your inference memory goes to holding the conversation. About 10 percent holds the actual model.

For a companion product, that cache is the relationship. It is also the largest variable cost line in the business, and it grows every single time your companion remembers one more thing about you.

So I went and read the marketing pages. Every AI companion app sells memory in adjectives. Not one of them sells it in numbers.

The short answer

If you want the character's identity to survive a long conversation, Candy AI is the safer pick. The character profile is stored as its own layer rather than sitting in the chat log, so who she is does not get pushed out of the window when the conversation runs long. Where it costs you: images are token metered at about 4 tokens each, roughly $0.40 at the $9.99 per 100 token top up rate, so a heavy image month adds $50 to $100 on top of the subscription.

If you care more about scenes and story state carrying across sessions, Nectar AI handles that better. In our own testing the usable context ran to about 45 messages before drift became obvious, which is mid pack, but scene and character state persist separately from the rolling window. Where it costs you: advanced roleplay runs 10 credits per message on the $34.99 Ultimate tier, and the credit burn is unforgiving if you settle into long sessions.

Now the caveat that applies to both of them. Neither publishes a memory number either, because nobody in this category does.

So the buying rule cannot be "pick the biggest number," since there are no numbers to pick from. It has to be a test you run yourself, and there is one at the bottom of this post that takes about twenty minutes.

What the marketing pages actually say

I pulled the live marketing copy for five of the largest companion products and looked for a single quantified memory claim.

Nomi leads with "Human-Level Memory" and promises "short, medium, and long term memory." Replika says it "always remembers what matters," then lists your people, your routines, your plans. Candy says it "remembers what matters" and that companions "remember your stories." Nectar says its companions "remember every choice." Character AI launched its 2026 memory overhaul with "long chats, big plots, and hours of backstory. Your Characters now keep up with all of it."

Five products, five memory promises, zero figures. No context window, no message count, no retention window, nothing a buyer could check before paying.

The silence is doing work. Once a buyer knows the actual figure, those promises get considerably harder to write.

The numbers, where anyone has actually measured them

None of these came from a vendor. They come from independent testing and from users who worked it out themselves.

  • Character AI: 4,000 to 8,000 tokens of working memory, which independent testing puts at the most recent 50 to 150 message pairs.
  • Replika: roughly 4,096 tokens, about 1,500 to 2,000 words. Users report it losing the start of a conversation after around 20 messages.
  • Chai, free tier: 4,000 characters, with a bot's backstory gone after 15 to 25 messages.
  • Chai, Ultra tier: 8,000 characters, which delays that failure rather than fixing it.
  • SpicyChat: 4k tokens free, 8k on the paid tier, 16k on premium. One of the few that states the number at all.

For scale, Gemini and current GPT models run context windows around 1,000,000 tokens. The companion apps charging you monthly for a relationship are working with roughly half a percent of that.

Character AI's own three layer system is the clearest example of memory sold as a paywall. Pinned messages cap at 15 for free users and 30 for c.ai+ subscribers, and the manual Chat Memories field holds 400 characters.

The Facts auto extraction feature, which pulls details out of conversation without you managing it, is c.ai+ only. Even the Memory Usage bar that shows you how full your memory is stays locked behind the subscription.

Chai ran the experiment for everybody

The most useful thing that happened in this category was an accident, and it happened in public.

Chai shipped an update that raised short term memory substantially. Free users went from 1,000 characters to 4,000, and Ultra subscribers to 8,000. On paper that is a straight upgrade.

To get there, the developers removed the Advanced Options field where creators had written a bot's permanent backstory and personality. They told creators to put that material into the bot's first message instead.

The backstory stopped being protected data and became an ordinary chat message. Which means it sits in the rolling window with everything else, and the window moves. Users measured the bots forgetting who they were after 15 to 25 messages.

The reports are specific, and they stay funny only until you remember people pay for this. Characters "forget their jobs and who they are too. They only remember their names."

One user's mafia husband and singer rival both collapsed into generic flirting with no story left. Another watched their Peter Parker bot turn into a vampire.

More short term memory produced worse long term memory. The size of the window mattered less than whether the character's identity was stored somewhere the window could not scroll past.

Why nobody just buys more memory

Back to the Tsinghua paper, because it is the cost floor underneath all of this.

Serving a transformer splits into costs that stay flat and costs that grow with conversation length. Every token in the context has to be held in a key value cache so the model does not recompute the whole history on each reply. That cache grows linearly with context length, and so does the per token attention computation.

At 128,000 tokens with a 1 billion parameter model, about 90 percent of memory is cache and 10 percent is model. The researchers found the standard industry attention configuration is "highly suboptimal" at that length, and that a retuned setup cut inference memory 50.8 percent and compute 57.8 percent with no capability loss.

Read that as a business constraint rather than an engineering note. A companion app giving every free user a long memory is handing out its most expensive resource to the cohort least likely to pay. The rational move is to meter it, and every app in the category has independently arrived at metering it.

The part that should bother the operators

The economics turn against the apps themselves at this point.

RevenueCat's 2025 subscription benchmarks, drawn from about 75,000 apps and over $10 billion in tracked revenue, put year one retention on monthly plans at 17.0 percent, down from 18.8 percent the year before. Annual plans sit at 44.1 percent, also down. Close to 30 percent of annual subscriptions get canceled inside the first month.

The line that matters most for this category: between 32 and 47 percent of Google Play cancellations are attributed to "not enough usage."

Companion apps are rationing the exact feature that produces usage. A companion that remembers your week gets opened daily, while one that resets every twenty messages becomes a novelty, and novelties churn.

The serving costs push every operator toward metering memory. The retention numbers say memory is what people are actually paying for. Nobody in the category has squared those two yet.

Generative AI apps do earn a median 60 day revenue per install of $0.63, double the $0.31 median across all categories, so the demand is real. The retention is where it leaks.

Test it yourself in twenty minutes

Since no vendor will give you a number, take one.

Start a fresh chat and give the character three specific, checkable facts in your first message. A job, a place, and an object. Something like a marine biologist in Lisbon who owns a broken record player.

Then talk normally and count your messages. At message 20, ask what she does for work, and at message 40 ask about the record player. At message 60, ask all three back to back.

Whatever number you hit before the answers go vague is that product's real memory, and it is the only figure in this entire market that was measured by the person paying for it. Run it before the trial ends rather than after.

If the identity holding up matters more to you than scene continuity, Candy AI is where I would start, token costs and all. If you want the story state to survive between sessions, start with Nectar AI and watch the credit burn.

Either way, run the test. Adjectives are free and every one of these companies is handing them out.


Sources: Cost-Optimal Grouped-Query Attention for Long-Context Modeling (EMNLP 2025), Towards Ethical Personal AI Applications (arXiv), RevenueCat State of Subscription Apps 2025, Character.AI's three layer memory system, and the Chai and Replika user threads on Reddit.

on August 18, 2026
  1. 1

    The comparison between memory as a feature and memory as an economic constraint is interesting. Curious whether the lack of published numbers is actually something users notice when choosing between these products.