Metlivi Blog

How to Evaluate an AI Companion’s Privacy, Memory, Controls, and Creative Skills

If you are deciding whether to try an AI companion, check four things separately: what the service says it collects, what memory you can inspect or change, whether you can remove chat history, and how it performs on a small task you care about. Product policies document stated practices; they do not independently verify how every step works. A remembered detail or a compelling conversation is evidence of your own experience, not proof of reliable memory or general capability.

September 30, 20266 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

What can you establish about data collection?

Start with the current privacy policy for the specific app and platform you plan to use. Look for categories of information, purposes, sharing with service providers, retention periods, and controls. “Your chats are private” is too broad to answer practical questions: a service may process chat content to generate replies while describing separate limits on advertising use or provider retention.

For one concrete example, Replika’s policy, last updated May 27, 2026, lists account and profile information, messages and uploaded content, interests and preferences, device and network data, and usage data. It says messages and preferences support personalized functionality, and that conversation data may be sent to third-party language-model providers to generate replies. The policy also describes using anonymized portions of messages for internal safety and performance work. These statements describe Replika’s published practices; they should not be generalized to other companions or treated as an independent technical audit. Replika Privacy Policy

Make a short checklist from the policy: Does the service receive text, voice, images, or all three? Which data is used for personalization? Are outside providers involved? How long are different categories kept? Can you request a copy or deletion, and what remains in the account after you delete a chat or account? If the policy leaves one of these unanswered, mark it unknown rather than inferring the answer from the app’s friendly tone or marketing.

Section 2

Does “memory” mean the same thing as reliable recall?

A memory feature can be inspectable and still make mistakes. A vendor’s explanation tells you how it presents the feature, not its success rate across ordinary use. Replika’s help page says its memory system supports personalization and lets users manually add details in a Memory tab. Its help center also says users can view, edit, add, and delete remembered details. Those are useful controls to verify in the app, but they do not establish that every detail will be recalled correctly in later chats. How does Replika’s memory work?

Research on personalized AI memory helps explain why users should test this directly. A 2025 CHI extended abstract reports interviews with six regular users of personalized AI tools and analysis of 54 public discussion threads. It describes differing expectations about what systems should remember and practices people use to organize or manage memory. The small interview sample and public posts can illuminate expectations and reported experiences; they cannot tell you how often a particular companion forgets or misremembers details. Users’ Expectations and Practices with Agent Memory

Try a simple, low-stakes check over several sessions. Give the system two or three distinct preferences relevant to a task, such as preferred tone, a fictional character’s name, and a planned change. Later, ask a natural question that would use one detail; then correct an error and see whether it applies the correction. Record whether it recalled the fact, confused it with another, used an outdated version, or asked you to repeat it. This is a personal spot check, not a benchmark, but it turns a vague impression into observable behavior.

Section 3

Which controls matter in practice?

Look for separate controls for stored memories, chat history, account deletion, and permissions such as microphone or photo access. These may do different things. Deleting a saved memory may not delete the conversation where you mentioned it; deleting visible history may not be equivalent to closing an account. Read the service’s instructions before relying on a control, and check the result in the interface afterward.

Replika’s help center makes this distinction concrete: its current article says the service does not offer a way to delete the entire message history without deleting the account, while the Memory section can be used to view, update, or delete selected remembered details. That is a product-specific description and could change, so check the current instructions for the app you use. Can I delete my conversations?

A useful control test is to add one harmless preference, confirm it appears in the memory view, edit or remove it, and then ask a later question that would have used it. Separately find the conversation-history and account-deletion instructions without executing an irreversible deletion. Note what each control claims to affect. This reveals whether the available controls match your own expectations before you share more.

Section 4

What does evidence say about ordinary creative tasks?

Published results exist for general AI writing assistance, but they should not be mistaken for direct evaluations of every companion app. In a 2025 experiment, 225 UK university students completed a short creative-writing exercise with or without ChatGPT. Participants using ChatGPT produced stronger results on the study’s creativity and writing measures, while reporting less enjoyment and perceived value for the task. The researchers measured a tightly defined activity: inventing uses for a hypothetical printer, under a ten-minute limit. That finding supports a narrow claim about that tool, sample, and task; it does not establish how a companion will help with your story, playlist description, invitation, or other personal project. “If ChatGPT can do it, where is my creativity?”

For a useful trial, choose a small task with criteria you can judge: ask for five names for a fictional café, a short scene in a specified tone, or three alternative endings to a story idea. Check whether it follows the constraints, offers genuinely distinct options, keeps details consistent, and improves after one correction. Save the prompt and output if you want to compare another session or product. This gives you evidence about that version and task, rather than a general verdict that it is “creative.”

Section 5

How to separate evidence, reports, and open questions

Treat product documentation as evidence of what the company currently says. Treat controlled studies as evidence about the particular participants, setup, product, and measures they tested. Treat user posts as reports that may reveal issues worth checking, not estimates of how common those issues are. In a study of Replika-related privacy discussions, researchers analyzed 111 Reddit posts; that design can document the kinds of information users said they shared, but a forum sample cannot represent all users or independently establish backend data handling. “I tell him everything that I do”: An investigation of privacy and safety implications of AI companion usage

For your decision, use a small evidence ledger: write down the policy statement, the control you verified, the result of your memory check, and the output of your creative task. Label each entry as “documented,” “observed in my trial,” or “not answered.” This keeps a persuasive interaction from standing in for a privacy answer, and keeps a published general-purpose writing result from standing in for testing the companion you might actually use.

Related reading

Keep exploring this topic