Metlivi Blog

What AI Simulated Readers Can—and Cannot—Tell You About Your Writing

If you are revising a page, guide, or set of instructions, a language model can help you spot wording that may confuse a reader. It can point to an unclear phrase, describe a possible misreading, or check whether a draft appears to answer a stated question. But a simulated reaction cannot show that your intended readers understood the text, brought the right context to it, or completed the task it describes. For that, watch a few real people from your audience try the task and explain what they understood.

September 30, 20266 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

What a model can usefully critique in a draft

A model can examine the words and structure you give it. Ask it to identify terms that might be unfamiliar, steps that seem to be missing, instructions with more than one plausible meaning, or a question the draft leaves unanswered. You can also ask it to state what evidence in the text led to each observation. This is useful as a first-pass editorial check: it creates candidates for things to investigate, not findings about actual readers.

Keep the task concrete. For example, with a page that explains how to choose a day trip, ask the model to point out which details a reader would need to compare options and where those details appear. Ask for specific passages and a possible misunderstanding, then check each suggestion against your purpose and the text. A fluent answer is not proof that any person would read the passage that way.

Persona prompts may add some useful structure, but a description such as “a busy first-time visitor” does not make the model that visitor. Research on persona prompting found that its benefits depended on whether the persona variables corresponded to patterns in the human annotations; across many subjective datasets, those variables explained little of the variation. Treat persona-based reactions as hypotheses, not as a representative sample of audience opinion. Hu and Collier, “Quantifying the Persona Effect in LLM Simulations”

Section 2

What simulated reactions cannot establish

A model response cannot establish what a particular reader noticed, inferred, remembered, or did. It receives text and a prompt, then generates an answer. A real reader encounters the material with their own knowledge, goals, distractions, expectations, and situation. Those factors shape whether a sentence makes sense and whether the person can act on it.

This difference matters most when the question is behavioral: Can someone find the right information? Do they interpret a direction as intended? Can they finish the task without a hint? A model can reason about these questions, but its account of what it would do remains generated text. It has not independently navigated the page or demonstrated that a person in your audience can use it.

Research into model simulations also warns against reading vivid or diverse responses as evidence of fidelity. A 2025 study of survey-response simulation found that the tested approaches struggled to accurately reproduce user responses; the models could retain viewpoints across demographic changes and had difficulty adapting to nuanced differences between profiles. That study does not directly test every writing critique or usability task, but it supports a practical boundary: a plausible-sounding simulated reader is not validation of how a real audience will respond. Yu et al., “An Analysis of Large Language Models for Simulating User Responses in Surveys”

Section 3

What a small real-reader test can reveal

A small qualitative test can show where actual participants hesitate, what they overlook, how they interpret a phrase, and whether they reach the intended endpoint. You can give each person a realistic task, let them use the draft or page, and observe what happens. Follow-up questions can clarify what they thought a word meant or why they chose a particular step. That combines observable behavior with the participant’s own explanation.

For a content task, define success before the session. If the text is a day-trip guide, for instance, you might ask a participant to choose an option that fits a stated preference and explain what information shaped that choice. Note whether they find the relevant details, which words or sections slow them down, and whether their explanation matches the distinctions the guide intended to convey. The example is a proposed test design, not a claim about any particular guide or outcome.

Government guidance on qualitative usability testing recommends recruiting people who could be users, assigning them a task, avoiding excessive instruction, and reviewing task completion and common errors afterward. Nielsen Norman Group similarly describes usability studies as a way to investigate whether content is easy to find and understand or whether people can complete a task. Both emphasize observing what participants do; that is evidence a simulated response cannot supply. GOV.UK, “Usability testing: qualitative studies”, Nielsen Norman Group, “Checklist for Planning Usability Studies”

Section 4

How to run a useful small test

Start with one decision or task the text is meant to support. Recruit people who resemble the intended audience in relevant ways, especially in their experience with the subject. Then write a scenario that gives them a reason to use the material without telling them where the answer is or which words to look for. A task that repeats a label from the page can inadvertently test word matching instead of whether the content helps someone accomplish their goal; NN/G recommends writing concrete tasks without clues that prime behavior.

During the session, let participants proceed with minimal guidance. Record what they do, where they pause, what they say in their own words, and whether they complete the task. If they ask for help, note the point at which they needed it. Afterward, ask them to describe what they understood and what they would do next. GOV.UK guidance suggests five to six participants for qualitative usability testing; NN/G generally recommends five for traditional qualitative studies. These are practical starting points for finding issues, not sample sizes that make findings representative of a whole population. If you need percentages or comparisons that generalize, you need a quantitative design with a larger sample and controlled measures. GOV.UK, “Usability testing: qualitative studies”, NN/G, “Quantitative vs. Qualitative Usability Testing”

Section 5

Turn observations into revisions, not broad claims

After a few sessions, look for recurring obstacles and consequential single failures. If several participants misunderstand the same phrase, that is a strong candidate for revision. If one person cannot complete a task, inspect the conditions and ask whether that failure matters for the intended audience; do not treat one session as a population-wide rate. A small qualitative study is meant to reveal problems and guide changes, not estimate how many people in a wider audience will have them. NN/G notes that qualitative findings depend on the participant sample and the researcher’s interpretation, while numeric usability claims require an appropriately sized quantitative study.

Revise the text, then test the revised version with fresh readers when possible. A model can help you explore alternate phrasings or flag new ambiguities between rounds. Keep the evidence types distinct in your notes: “the model predicted this might be confusing” is a prompt for investigation; “two participants interpreted this step differently and one did not reach the endpoint” is an observation from a test. Neither alone proves that every reader will have the same experience, but the second tells you what happened in the task you observed.

Section 6

A practical division of work

Use the model to generate questions about the draft. Use real readers to answer questions about actual comprehension and task performance. When the task depends on a specific audience, environment, or prior experience, recruit for those conditions rather than asking a model to invent them. And when you need a reliable rate or comparison, design a quantitative study instead of stretching a handful of qualitative sessions into a statistical claim.

This division lets simulated critique speed up revision without confusing a plausible reaction with evidence of reader success. The decisive question is not whether the model can produce a convincing reader voice; it is whether the intended reader can understand the material and do what the text is there to help them do.

Related reading

Keep exploring this topic