Metlivi Blog

How to Observe an AI Game Dialogue Playtest in 20 Minutes

For a small studio testing one generative dialogue scene, observe what players do with the NPC’s answers: whether they try different questions, act on useful information, recover from invented facts, and leave dialogue to finish the scene. A short questionnaire can capture what players say they felt or understood afterward; it cannot, by itself, show the choices and detours that happened moment by moment. Use a simple event sheet during play, then ask focused follow-up questions. Treat the observations as evidence about this scene and build—not as a measure of satisfaction or proof of how all players will behave.

September 27, 20266 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

What should you observe in a generative dialogue scene?

Choose a scene with one clear objective and an NPC whose answers could change how the player pursues it. For example, in a fictional scene, the player must open a sealed greenhouse gate before a ventilation cycle ends. A maintenance NPC can offer clues through free-form conversation. The intended route is to find a blue valve handle in the tool shed, but the NPC may occasionally invent a detail, such as claiming the handle is in the flooded pump room.

The four indicators below focus on player behavior and the scene’s consequences. They do not require judging whether a question is clever or whether a player appears to enjoy the game.

Indicator: **Varied questions** — Record an event when…: The player changes the wording, subject, or approach after an answer—for example, asks where a valve is, then asks what color it is or which room is safe. — State evidence to note: What they asked, what the NPC answered, and whether the next question followed from that answer. Distinguish a genuinely different inquiry from a near-repeat. — Ask afterward…: “What were you trying to learn with those questions?”

Indicator: **Answer changes the next action** — Record an event when…: The player moves, inspects something, or changes their plan in a way that follows the NPC’s answer. — State evidence to note: The answer, the player’s next action, and any visible alternative they bypassed. Mark the connection as clear, plausible, or uncertain; an action after an answer does not prove the answer caused it. — Ask afterward…: “Which part of the conversation, if any, affected what you did next?”

Indicator: **Invented fact causes a wrong turn** — Record an event when…: The player follows a specific unsupported or false NPC claim and takes an unproductive action because of it. — State evidence to note: The exact claim, the action it prompted, the cost or detour visible in the build, and how the player discovered or corrected the error. Check the scene’s intended facts after the session before labeling a claim invented. — Ask afterward…: “What made that route seem like the right one? When did you decide to change course?”

Indicator: **Player can stop dialogue and complete the scene** — Record an event when…: The player exits, pauses, or declines further conversation and can still pursue and complete the objective. — State evidence to note: Whether an exit is visible, what happens after leaving, whether useful progress remains possible, and whether the player reaches the scene’s completion state. Record blockers separately from a player choosing to keep talking. — Ask afterward…: “Did you feel able to leave the conversation? What made you continue or stop?”

These are event definitions, not scores for quality. Note the player’s actual words and actions rather than inferring intent from facial expression, silence, or playtime. If an event does not occur, record “not observed in this session,” not a presumed failure.

Section 2

Keep the observation sheet tied to the build

Before each session, record the build version, starting objective, known scene facts and allowed completion state. During play, write down the player question, the NPC answer, the next visible action and any committed game-state change. A line that sounds helpful can still be wrong if it points to a location that does not exist in this build.

After the session, compare suspected invented facts against the actual scene sheet. Mark an unsupported claim as confirmed only when the scene facts contradict it; otherwise label it unresolved and investigate. Keep a separate note for facilitator prompts or technical interruptions, because those can change what the player does next. This small evidence trail makes the later discussion more precise without turning one playthrough into a universal result.

Section 3

How to run a bounded 20-minute session

Tell the participant: “Please play this scene as you normally would. You can talk to the character or leave the conversation whenever you want. I may stay quiet so I can watch what you try.” Avoid teaching them to ask varied questions or warning them about invented facts; that would change the behavior the session is meant to reveal. If they ask for help, respond consistently, record the intervention, and treat later behavior with that help in mind.

A workable schedule is:

**Minutes 0–2: Set up.** Explain the task and controls without describing the intended solution. Start the clock when the player gains control, and start the scene in the same state for each participant.

**Minutes 2–15: Observe.** Let the player explore and talk. Log each relevant answer and the next action, especially when a claim sends them somewhere. Keep neutral prompts such as “What are you thinking?” for moments when the player has stopped and needs a prompt to continue; write down that you intervened.

**Minutes 15–20: Close and ask.** If the player has not finished, stop at 15 minutes of play and record the current state rather than implying they failed a speed test. Ask the follow-up questions tied to observed events, then ask one broad question: “What, if anything, was unclear about the conversation or your next step?” Record self-report separately from observed behavior.

The 20-minute limit is a practical session boundary for this example, not an empirical benchmark. A player who spends longer talking has not thereby demonstrated greater satisfaction, and a player who finishes quickly has not necessarily understood or enjoyed the scene. If the task requires more time in your build, adjust the schedule before sessions and keep it consistent.

Section 4

Separate what happened from what the player says

In [The Turing Test postmortem](https://www.gamedeveloper.com/business/postmortem-building-i-the-turing-test-i-around-a-secret-mechanic), design director David Jones describes testing puzzles with students, collecting fun and difficulty ratings and playtime, and placing greater importance on watching player behavior. He describes combining ratings and observation to assess the difficulty curve. For a dialogue scene, the useful lesson is to keep the two kinds of evidence together but distinct: the event log shows what the player did; the follow-up records their explanation or rating. Neither automatically explains the other.

Klei’s [Mark of the Ninja postmortem](https://www.gamedeveloper.com/design/classic-postmortem-klei-entertainment-s-i-mark-of-the-ninja-i-) describes frequent tests with fresh players to examine design assumptions. The team looked for the motivation behind complaints and watched where new players struggled, then adjusted cues and design. Applied here, a player’s wrong turn is a clue to investigate—not a prompt to simply make the NPC’s dialogue longer. Ask what in the answer, interface, or scene made the route persuasive, and check the recording or build state before deciding what to change.

Ubisoft’s [Teammates announcement](https://staticctf.ubisoft.com/8aefmxkxpxwl/2QCAorjku7w7gH1LGORV3t/6e8f347be3ecab7daa4769e5300086bc/Ubisoft_Unveils_%C3%A2__Teammates%C3%A2____Its_First_Playable_Generative_AI_Experience_Through_Closed_Player_Testing.pdf) describes a closed player test for a playable generative AI experience. It establishes that the team announced closed testing; it does not provide published player outcomes in the cited announcement, so it cannot support claims about what players did or whether the experience succeeded.

Section 5

Turn observations into a next build decision

After the session, review the timeline and connect each NPC answer to the next player action. For every apparent wrong turn, verify the relevant scene facts, then identify the likely source: the NPC supplied a false detail, the player misread a true answer, or a location or interaction cue pointed them elsewhere. These explanations are hypotheses until checked against the recorded sequence and, where useful, another session.

Use the four indicators to pick a concrete follow-up change or test. If the player asks several distinct questions but gets answers that do not guide action, review whether the scene offers actionable information. If a specific invented claim sends them to the wrong room, consider whether the scene needs a way to verify the claim or recover without a dead end. If the player cannot leave dialogue and finish, inspect the exit and objective flow. If they do leave and complete the scene, note that the path was available in this build; ask what they understood before concluding it was clear.

A single 20-minute session can expose a particular confusion or missing route, but it cannot establish how common that issue is. Keep the observer’s event notes, the verified scene facts, and the player’s retrospective answers separate when deciding what to investigate next. That gives a small team a useful account of how this player navigated this generative dialogue scene, beyond a questionnaire alone.

Related reading

Keep exploring this topic