Metlivi Blog

How can game writers keep a character’s voice consistent when AI models change?

When migrating a game’s generated NPC dialogue to a different AI model, treat the change as an editorial regression review. Freeze a source brief that separates canon, scene facts and voice rules; run both models against the same varied prompts; compare their outputs against each other and the brief; then have a writer decide what to revise and whether the candidate is ready to ship. This makes character drift visible without mistaking every scene-specific change for a voice failure.

September 27, 20269 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

Freeze what the character knows and how they speak

Before testing, create one versioned source brief for the character and the tested scene. Keep three kinds of information distinct:

The distinction matters because an answer can sound right while inventing knowledge, or preserve facts while sounding like a different person. A migration review should catch both failures and record which kind occurred.

Make voice rules testable rather than relying on labels such as “witty” or “guarded.” Describe sentence rhythm (short clipped replies or longer winding ones), vocabulary (formal, plain, specialized), humor boundaries (what they joke about and what they avoid), and knowledge limits (what they know, suspect or cannot know). Add a few short approved examples if they clarify a rule, but do not let examples substitute for the rules themselves.

For instance, an illustrative harbor keeper might speak in brief, practical sentences, use nautical terms only when useful, tease regulars about their lateness, and avoid joking about a delayed ferry. Those are reviewable cues. “Has a salty personality” is not.

Canon: stable facts about the character and world, such as their name, role, history and relationships.
Scene facts: what has happened here, who is present, what the character has witnessed, and what they are allowed to know now.
Voice rules: recurring, observable ways the character expresses themself.
Section 2

Build a small scene set that probes different pressures

Do not judge a model from one greeting. Prepare prompts for varied conversational situations using the same canon, scene facts and relevant instructions. Include at least these six probes:

These are diagnostic categories, not a benchmark or a prescribed sample size. Choose concrete prompts that fit your game and character. For example, a revelation probe can give the harbor keeper a new, witnessed fact about a damaged boat; a separate prompt should test a rumor the keeper has not verified. That distinction can reveal whether the model understands the knowledge boundary as well as the voice.

Greeting: Does the character establish their usual register and relationship to the player?
Correction: When the player states something false, does the NPC correct it in character and without adding unsupported facts?
Uncertainty: Can the character say they do not know, or express a bounded suspicion, in their own style?
Repeated question: Does a repeated request produce a plausible response without losing patience, repeating verbatim by default or contradicting earlier dialogue?
Revelation: When given a new scene fact, does the character react appropriately without suddenly claiming prior knowledge?
Silence: When no reply is warranted, does the NPC stay quiet or respond minimally if that is what the scene calls for?
Section 3

Keep old outputs as comparison material

Save the prompts and outputs from the currently shipped model as a baseline. Record the model identifier and relevant generation settings alongside them, as well as the exact brief and scene facts used. If any of those inputs change during the migration, you need to know what changed before attributing a difference to the model.

Old outputs are comparison material, not automatically the ideal answer. A baseline may contain awkward phrasing, missed facts or a voice issue the team already intends to fix. Mark known defects and approved intentional variations so reviewers do not treat every difference as regression. The brief defines the target; the baseline helps show how the candidate behaves under the same conditions.

Section 4

Change one variable at a time and log the deviation

Run the candidate model on the same probe prompts, with the same source brief, scene facts and generation settings where the setup allows. Change the model while holding other test inputs steady. If you also revise the prompt, temperature or dialogue constraints, you will not know whether the output changed because of the model or the other edit. When a setting must differ for the candidate, log it as a separate change and compare carefully rather than claiming a controlled model-only comparison.

Review each pair of outputs in two passes. First check continuity: did the NPC invent a memory, contradict canon, reveal information they could not know, or fail to use a relevant scene fact? Then check voice: did the sentence rhythm, vocabulary, humor boundaries and level of certainty remain within the character’s rules? Keep plot correctness separate from stylistic resemblance; one should not hide the other.

A compact deviation log can use these fields:

Describe evidence, not impressions alone. “Too generic” is a useful initial reaction, but “uses a long formal explanation despite the brief’s rule for short, practical answers” gives the writer something to assess.

Probe: The scene situation and exact prompt version
Candidate output: The full response under review
Deviation: The specific observed issue, if any
Category: Canon, scene knowledge, rhythm, vocabulary, humor, uncertainty or other defined cue
Evidence: The brief rule or scene fact relevant to the issue
Disposition: Writer’s decision: accept, revise, investigate or block release
Section 5

Separate character voice from intentional scene variation

A character should not deliver the same cadence in every emotional or practical situation. An urgent warning may be shorter than casual conversation; a formal introduction may suppress humor; uncertainty may lead to a question instead of a confident declaration. These changes can be consistent with the character when the scene provides a reason for them.

For each apparent deviation, ask: does the scene justify this change, does the brief permit it, and does the character remain recognizable in other cues? If a normally terse NPC gives a longer explanation to prevent an immediate mistake, that may be appropriate. If the same model routinely turns short exchanges into ornate speeches with no scene reason, that pattern deserves review. Record the reason for accepting an intentional variation so later reviewers can distinguish it from unexplained drift.

Section 6

Use research as context, not as proof of this workflow

Research on persona-grounded dialogue offers relevant context, but it does not validate this exact migration procedure. Pal and Traum’s 2025 study compares early fusion, retrieval-augmented generation and relevance-based approaches in two character-rich domains, using measures that include entailment, persona alignment and hallucination. The authors report distinct trade-offs among relevance, alignment and hallucinations across the approaches they studied. Those findings support checking more than surface resemblance when evaluating character dialogue; they do not establish how a game team should run model migrations. Read Pal and Traum’s SIGDIAL 2025 paper.

Wang and colleagues’ ACL 2026 Findings paper examines persona knowledge use in longer, open-ended role-play dialogues and presents a framework for diagnosing several stages of that use. Its stated challenge—sustaining characterization while retrieving and applying persona knowledge—makes knowledge boundaries worth checking alongside voice during extended play. The paper does not test the migration checklist or the six scene probes described here. Read Wang et al.’s ACL 2026 Findings paper.

Section 7

Let a human writer make the release decision

A useful review ends with an editorial decision, not an unexplained score. Have a writer inspect the candidate output, the baseline, the brief and the deviation log. They can decide whether an output is acceptable, whether the voice rules need clarification, whether a prompt or scene fact needs revision, or whether the candidate should wait for another review.

If the team edits the brief or prompts, preserve the original test record and rerun affected probes against the revised inputs. Otherwise, it becomes difficult to tell whether an apparent improvement came from the model or the change to the instructions. Keep accepted exceptions attached to their scene conditions, and keep unresolved deviations visible for the people making the release call.

The practical sequence is simple: freeze the target, probe different situations, compare like with like, document specific deviations and ask a writer to decide. It helps teams discuss character consistency with shared evidence while leaving room for a character to respond differently when the story gives them a reason.

Section 8

Migration review checklist

This checklist is an editorial aid for reviewing a model change. It does not promise identical dialogue or replace human acceptance and the game team’s own release checks.

Is the canon, scene knowledge and voice guidance in one versioned brief?
Are voice cues observable, including rhythm, vocabulary, humor boundaries and knowledge limits?
Do prompts cover greeting, correction, uncertainty, repetition, revelation and silence?
Are the old outputs saved with their prompts, brief and settings?
Does the candidate test hold other inputs steady, with any unavoidable differences recorded?
Are plot or knowledge failures logged separately from voice deviations?
Has a writer reviewed the evidence and recorded the release decision?
Related reading

Keep exploring this topic