Metlivi Blog

Look for the evidence trail behind an AI companion feature

You cannot prove from a product page alone that an AI companion feature was thoroughly tested, but you can identify whether the visible evidence is strong enough for the decision you face. Start with six questions: Which exact version was tested? What intended use and exclusions were stated? Which realistic scenarios were covered? What failures and recovery paths were observed? Was any review separated from the feature’s builders? What happens after release when behavior changes? Missing answers do not automatically prove a bad feature; they reduce confidence and should narrow your use. Keep the first trial reversible, avoid sensitive disclosures, leave optional tools off, and do not pay for a broad promise that is supported only by a polished demonstration.

August 27, 20268 min readHome, Safety, Pets & Sustainable LivingBy Metlivi Editorial Team
Section 1

Rung one: identify the tested system

A model name is not enough. Look for the app version, model or service version, enabled tools, memory setting, language, platform, date, and paid tier used in evaluation. A companion feature may change when the operator updates the model, prompt, retrieval source, moderation layer, voice pipeline, or tool permissions without changing its marketing name. Evidence that names none of these cannot be reliably matched to what you see today. Compare release notes, help pages, in-product labels, and the evaluation date. If the tested configuration is unclear, record “version not established” rather than assuming the newest screen inherited older results. This first rung prevents every later result from floating free of a specific product.

Section 2

Rung two: compare intended use with the promise

Good documentation explains what the feature is meant to do and where it should not be relied upon. Translate the advertisement into tasks: casual text exchange, activity suggestions, image response, voice input, web retrieval, reminders, or actions in connected services. Then check whether evaluation covers the same tasks. A text-only assessment says little about voice, images, long memory, external tools, or public interaction. The FTC complaint concerning an AI detector describes a promoted accuracy claim that was not tested across different conditions of use; the broader lesson is to match a claim to evidence rather than borrowing confidence from a narrower test. When scope differs, downgrade the claim to the proven task.

Section 3

Rung three: inspect scenarios and failure examples

A percentage without scenario definitions is difficult to interpret. Useful evidence describes ordinary, edge, and adversarial inputs; account state; language; modality; relevant user groups; and scoring rules. It also provides examples of what counted as failure, disagreement, refusal, or unresolved behavior. Look for interrupted connections, stale history, shared devices, ambiguous instructions, long conversations, permission denial, and tool errors where relevant. Perfectly selected demonstrations and average scores can hide rare but important failures. NIST’s AI resources emphasize test, evaluation, verification, and validation in context. Ask whether the scenarios resemble the way you will use the feature, and whether recovery was tested when the ideal path broke.

Section 4

Rung four and five: seek limits and review independence

Credible evidence makes limitations visible near the result. It distinguishes known gaps from items not tested and explains mitigation without implying that every risk disappeared. Check whether a group outside the feature’s immediate builders performed assurance work, whether outside specialists participated, or whether at least a held-out set was protected from tuning. Independence is not a badge of perfection; it reduces the chance that one team chooses both the questions and the flattering interpretation. OpenAI’s system cards illustrate the kind of trace a reader can inspect: model scope, evaluation stages, red-team work, observed risks, and product mitigations. Smaller products may publish less, but should still answer concrete questions about method and boundaries.

Section 5

Rung six: verify monitoring and change control

Testing ends at release only on paper. Real behavior changes with new versions, policies, languages, tools, and user patterns. Look for dated release notes, a channel for reporting reproducible issues, an incident or status route, clear version changes, and evidence that important scenarios are rerun. Confirm whether a major update changes permissions, memory, sharing, billing, or deletion. A feature with impressive launch evidence but no visible maintenance path becomes harder to assess over time. Conversely, a concise change log that names the altered surface, known limitation, and retest scope can be more informative than a permanent “tested” badge. Record three dates: current feature version, latest relevant evaluation, and your last low-disclosure check.

Section 6

Choose a use level from the evidence ladder

Score each rung as visible, partial, or absent, then select a reversible use level. With weak version and scope evidence, stay with neutral text and no optional tools. With credible scenario and recovery evidence, you may test the covered task while keeping unrelated permissions off. If payment, public posting, external actions, or persistent memory are involved, require stronger documentation before enabling them. Do not turn the ladder into a public ranking; it supports one decision for one configuration. Save links and dates, not screenshots containing other people’s material. Recheck after updates. The useful question is not “Is this feature universally safe?” but “Which exact use is supported by current evidence, and what remains outside it?”

Related questions

Common questions

Does missing documentation prove the feature was not tested?

No. It means an outside reader cannot verify scope, method, or result, so use should remain narrower until questions are answered.

Is one high benchmark score enough?

No. You need the tested configuration, task match, scenario distribution, scoring rules, failure examples, and product-level recovery.

What is the quickest first check?

Identify the exact version and date, then see whether the published scenarios match the feature and language you plan to use.

Related reading

Keep exploring this topic