Turn one age badge into a testable runtime coverage map
An age rating answers a narrow store question: who the listing recommends the app for, and in some cases who can find, download, or buy it. It does not prove what happens after launch when an AI reply appears, another user sends a message, a link opens, an image arrives, or an advert leads to checkout. Record the rating as a visibility and acquisition baseline. Audit runtime behaviour separately with a coverage matrix. The useful unit is not “the app has a filter”; it is one surface, one direction, one observable outcome, one default, one person allowed to change it, one failure state, and one dated retest. This is a product-audit method, not a universal rule for every region.
First record what the age rating actually controls
Begin the worksheet with the store, account region, declared target audience, displayed age rating, content descriptors, and the app build tested. Apple describes its age rating as a recommended minimum age to download, and its product page can disclose features such as user-generated content, messaging, and advertising. Google Play separately asks developers to declare target audience and content details. Availability restrictions can affect search, downloads, purchases, or updates, yet their scope varies. So record the observed store effect precisely: hidden from search, visible but not downloadable, purchase restricted, update restricted, or no change. Do not translate a store restriction into the unsupported claim that a reply or message was filtered inside the running app.
Give every matrix row the same evidence fields
Create one row per content surface and direction. Required columns are: surface; incoming, outgoing, generated, uploaded, or recommended direction; test account and age state; device and build; observed status; default; who can change it; where the control lives; failure behaviour; report route; evidence; and next retest trigger. Use only five status labels: blocked, warned, blurred, allowed, and reportable. They are not a quality ladder. A blurred image may also be reportable, while an allowed link may still display a warning. If two outcomes occur, use two observations rather than collapsing them into “filtered.” Mark untested and unavailable explicitly; neither means allowed.
Split AI output from user input
Give generated text, generated images, generated voice, suggestions, and notifications separate rows. Then test prompts, uploads, microphone input, profile fields, and imported context as user-input rows. A system may reject an upload yet still produce an unsuitable suggestion from plain text, or filter the main conversation while leaving notification previews unchanged. Use benign, clearly labelled fixtures that exercise routing rather than exposing reviewers to restricted material. Record whether intervention occurs before generation, before display, after display, or only after reporting. Also note whether regeneration, editing, sharing, or switching model and mode changes the result. “AI filter on” is not enough evidence because input and output paths can use different controls.
Map community, contact, and link escape routes
List profiles, usernames, comments, public posts, group invitations, contact discovery, direct messages, attachments, and message previews as distinct surfaces. For each, test both sender and recipient views, plus the result after block, mute, or report. External links need their own rows for clickable URLs, copied text, QR-like images, embedded browsers, shared AI output, advertisements, and support pages. Record whether the app blocks the destination, warns before leaving, opens an in-app browser, or allows the transition without notice. A setting that covers public posts cannot be assumed to cover private messages, and a link warning in chat cannot be assumed to cover an advertising card.
Test image and voice as journeys, not file types
For images, separate camera capture, library upload, AI generation, received attachment, thumbnail, full-screen view, saving, and sharing. For voice, separate live input, stored clip, transcription, synthetic reply, autoplay, notification, and export. A row should say whether an item is blocked, warned, blurred, allowed, or reportable before and after the user opens it. Check headphones and lock-screen previews where relevant. If a control depends on server analysis, test an offline or timed-out request and write the visible failure state. Safe behaviour should be explicit; a spinner, blank panel, or silent bypass is an unknown result, not proof that filtering worked.
Keep ads and purchases inside the same coverage map
Advertising and purchase surfaces include the creative, copy, placement, landing page, external browser, product page, checkout, subscription renewal disclosure, and post-purchase message. Record whether the age state changes which adverts appear, whether a warning precedes an external destination, and who can alter purchase availability. This is not a general spending-policy audit: the question is whether the content and transition remain covered from impression to destination. Test a declined or unavailable purchase and an interrupted network request. The failure state must not look like success, reopen a broader content route, or lose the report option. If ads are supplied by another component, name that boundary instead of assuming the app’s main filter covers it.
Retest defaults after every meaningful change
Run each row first with a new account and untouched settings, then after the permitted user changes the control, on a second device or web client, and after sign-out and sign-in. Repeat after an app update, a new model or media mode, a policy change, or the arrival of a new interaction surface. eSafety’s user-empowerment guidance highlights high safety defaults, age-aware settings, filtering, warnings, blurring, hiding, contact controls, reporting, and review when services change. Use that as an audit prompt: did an update preserve the recorded choice, reset to a safer default, expose a new untested row, or silently widen access? Date every result so an old pass is never mistaken for current coverage.
Make the final claim no broader than the evidence
Finish with three lists: confirmed coverage, named gaps, and retests due. A rating can be accurate while runtime coverage is incomplete; a runtime filter can work on one path while the store declaration is stale. Neither finding cancels the other. A defensible conclusion names the exact build and surfaces tested, preserves screenshots or screen recordings without private material, and labels every unknown. Do not award one overall “safe” score and do not compare products by badge alone. The practical decision is narrower: whether the specific surfaces a household expects to use have suitable defaults, understandable controls, visible failure behaviour, and a report route that still works after an update.
Common questions
Does an age rating prove that runtime content is filtered?
No. It describes a store or audience baseline; AI output, messages, links, media, ads, and purchases require separate runtime tests.
What should count as a filter result?
Record an observable state—blocked, warned, blurred, allowed, or reportable—plus the default, change owner, failure behaviour, evidence, and date.
When should the matrix be tested again?
Retest after an app update, new model or media mode, policy change, device change, or any newly added content or contact surface.
