Metlivi Blog

Can a Fixed Game Script Feel More Interesting With a Few Generated Dialogue Lines?

Yes—if generation adds local variation while authored beats still control what happens next. Keep the objective, turning points, and consequences fixed; let a small number of dialogue slots vary within explicit limits, such as a character’s tone or a detail they notice. Then compare that version with a scripted-only build using the same scenes and player choices. Look for dialogue that players describe as fresh or responsive, while checking whether either version creates more confusion about characters, events, or game state.

September 30, 20266 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

What the hybrid approach is meant to change

A hybrid narrative keeps its major events under direct authorial control and uses generation for selected lines around those events. For example, the scene may always begin with a shopkeeper closing up, present the same player choice, and end with the same clue. A generated line might change how the shopkeeper comments on the weather or greets the player, but it should not invent a new clue, change the choice, or imply that an earlier event happened differently.

That boundary matters because dialogue can feel varied without changing the story’s state. Interactive narrative research describes storylets as discrete narrative modules whose availability can depend on the current game state. Another interactive narrative system combined high-level story direction with autonomous character behavior, and adapted the storyline when player actions created inconsistencies. These examples support a useful design principle: treat authored events and current state as structural constraints, then allow variation only where it cannot contradict them. (Sketching a Map of the Storylets Design Space; Mixing Story and Simulation in Interactive Narrative)

Section 2

Pick slots where repetition is noticeable but the stakes are low

Start by identifying dialogue that repeats across visits or playthroughs and does not carry essential information. A greeting, a brief reaction to a player’s previous action, or a line of scene texture may be a candidate. A reveal, promise, instruction, or statement about an item’s ownership is a poor first slot: changing its meaning could disrupt the plot or players’ understanding of the world.

For each slot, write down its job before writing or generating variants. Record who is speaking, what the character knows, what just happened, the emotional range that fits, and which facts must remain unchanged. Set a clear length limit and define forbidden additions, such as new quests, items, relationships, or past events. A slot is bounded when a line can be replaced without altering the scene’s playable outcome or the facts players need to carry forward.

This is a design recommendation, not a claim that every generated line will feel fresh. Some players may not notice a small variation; others may find that even a well-formed line distracts from a carefully paced scene. The test should establish whether the variation serves this particular game and audience.

Section 3

Make the scripted-only version a fair comparison

Build two versions of the same short playable section. In the control version, use the existing authored line in each candidate slot. In the hybrid version, retain the same authored beats and substitute generated dialogue only in those slots. Keep the scene order, choices, rewards, character appearances, and play instructions the same. If those elements differ, it becomes harder to tell whether players reacted to dialogue variation or to some other change.

Decide what you want to learn before recruiting players. A practical question is: “Does bounded dialogue make this scene feel more responsive or less repetitive, without reducing understanding of what happened?” Use comparable players and give them the same task. If participants play both versions, consider that the first run may teach them the scene; if different groups play each version, differences between groups may affect the comparison. Note the limitation whichever setup you choose.

Game user research guidance recommends choosing methods based on the research objective and combining observations, surveys, and interviews when useful. Observation shows what players do, surveys can make ratings easier to compare, and interviews can help explain behavior, though each method has limits. For this test, record play behavior and collect a brief post-play rating or interview rather than relying on a single general question like “Was it fun?” (Choose the Right Playtest Method)

Section 4

Track creative variation and state consistency separately

Use two separate scorecards so a lively line cannot conceal a continuity problem. The first measures whether variation was perceptible and useful. After play, ask players whether the dialogue felt repetitive, whether character reactions seemed responsive, and whether any line stood out. Ask for a specific moment rather than only a numerical score. During play, note whether players pause, reread, or comment on a line, but do not assume that attention automatically means enjoyment.

The second scorecard checks consistency. Compare each generated line with a short list of state facts: character identity and knowledge, location, prior actions, current objective, and facts established earlier in the scene. Log contradictions, unsupported new claims, and moments when players appear unsure what happened. Also ask players to summarize the scene’s key event and next objective; a mismatch can point to a comprehension problem, though a small playtest cannot by itself establish why it occurred.

State tracking deserves its own attention. A study framing Dungeons & Dragons as a dialogue task treated generating the next turn and predicting game state from dialogue history as related but distinct challenges. Its dataset included partial state annotations, and the authors evaluated both generated dialogue and state prediction. That research does not establish that a particular hybrid design will work, but it reinforces why narrative variation and state accuracy should be measured separately. (Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence)

Section 5

Read the results as a design decision, not a verdict on generation

Compare the versions on the same measures. Did players notice the dialogue variation? Did they describe a specific reaction as more fitting or less repetitive? Did they understand the same scene facts and objective? Did the hybrid version introduce any continuity error? A useful result can be mixed: players may enjoy a varied greeting while finding generated exposition harder to follow. That suggests keeping the greeting slot and removing or narrowing the exposition slot.

Review the actual lines players saw alongside their comments and state-check notes. Averages can hide one serious contradiction, while memorable anecdotes can exaggerate a rare event. If a line breaks a required fact, treat it as a defect to fix before interpreting the enjoyment results. If players do not notice variation, check whether the slot is too brief or too subtle before expanding generation across the script.

Research on a role-playing game that compared static, rephrased, hybrid, and fully generated dialogue used behavioral logs and post-game surveys with 64 participants. Its reported results suggest that levels of generative agency can shape interaction and engagement, but those findings belong to that game and study. They are not a prediction for another title. The practical takeaway for a small design experiment is to compare dialogue approaches directly in the intended play context. (Quest of Aivengarde: Comparative Study of Player Experience Across LLM Dialogue Systems)

Section 6

Expand only the slots that pass both checks

If the hybrid version yields lines players notice and appreciate, while preserving the same scene understanding and state facts, try a second small set of low-stakes slots. Keep the fixed beats fixed and rerun the comparison. If a slot adds little perceived variation, or repeatedly creates confusion, return it to authored dialogue or tighten its boundaries.

A few carefully chosen generated lines can make repeated moments feel less identical, but whether they make a game more interesting is an empirical question for its players. Preserve the story’s essential events, vary only dialogue with a clear local purpose, and use a controlled playtest to judge both the creative payoff and the continuity cost.

Related reading

Keep exploring this topic