Metlivi Blog

How to Keep Branching Game Stories Testable

When a branching story grows, test the decisions and state changes that make paths behave differently—not every possible route as a separate end-to-end script. Model the narrative as a graph, define the conditions and outcomes at each decision, and build a small suite that covers critical transitions, rejoin points, and failures. Then use risk-based route sampling and human playtests to catch issues that structured checks cannot judge. This gives narrative designers and QA teams a repeatable way to find defects without claiming exhaustive path coverage.

September 27, 20267 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

Start with a graph that records behavior

Represent each playable passage or scene as a node and each choice as a directed edge. Annotate edges with their conditions and effects: for example, `has_key = true` enables “Unlock the gate,” which sets `gate_open = true`. Mark endings, loops, and rejoin points explicitly. A rejoin point is where distinct routes meet again; it is a useful place to check that the story can continue from multiple histories without carrying unintended state.

The graph should reflect what the game actually evaluates, not only the prose structure. Record which variables a choice reads, which it writes, and which later nodes depend on them. Include defaults, reset rules, and any one-time effects. If a flag is set in one branch and never cleared, that may be intended; documenting it makes the consequence visible for review and testing.

This structure also helps locate decision points that are easy to miss in a long script. A 2024 paper by Alexey Tikhonov studies detection of character decision points in branching narratives and proposes a dataset based on Choose Your Own Adventure game graphs. Its task concerns identifying narrative decision points, not validating a QA method; it can inform how teams inventory choices, but it does not show that the test approach here is effective. [Tikhonov, “Branching Narratives: Character Decision Points Detection”](https://aclanthology.org/2024.games-1.8/)

Section 2

Test state transitions, not just scene visits

A test that confirms a node appeared can miss a broken choice. For each important choice, check three things: whether the choice is available under the intended condition, whether selecting it applies the expected state change, and whether the next node and visible outcome match that state. These checks treat the story as a state-transition system: given a starting state and an action, verify the resulting state and destination.

For example, a test for choosing “Show the map” might assert that the map is shown, `trust` remains unchanged, and the route reaches the shared observatory scene. A paired test begins with `has_map = false` and asserts that this choice is unavailable or follows the specified alternative. The exact expected behavior depends on the narrative specification; the important point is to assert it explicitly rather than infer correctness from a passage title.

At rejoin points, test more than arrival. Compare the state each route is meant to preserve, change, or discard. A guard who was persuaded on one branch might remain an ally after the reunion, while a temporary disguise should expire. Make those rules part of the expected post-rejoin state. If routes are meant to converge completely, assert the shared state; if they should retain meaningful differences, assert those differences too.

Section 3

Use a fictional example to make coverage visible

Suppose a short mystery has a decision at the archive. The player can request help, sneak in, or use a borrowed key; each route reaches the same corridor, then a later choice determines whether the player takes a sealed letter. The following fictional matrix tracks a compact set of test obligations. “Covered” means a test is planned for the specific obligation, not that the whole route or every combination has been tested.

**A:** `trust = high`; ask the archivist for help. Help option appears; `trust` stays high; route reaches corridor. Choice availability and transition

**B:** `trust = low`; ask for help. Help option is hidden or declines, as specified. Negative condition

**C:** `has_key = true`; unlock side door. Door opens; route reaches corridor; key is consumed only if specified. State effect and rejoin

**D:** `has_key = false`; attempt side door. Door cannot be opened; no success flag is set. Negative assertion

**E:** From corridor, take sealed letter. `has_letter = true`; later evidence scene offers the letter-specific line. Downstream outcome

**F:** From corridor, leave letter. `has_letter = false`; letter-specific line is absent. Outcome contrast and negative assertion

This is a decision aid, not a coverage percentage or a universal minimum suite. It makes omissions legible: here, the low-trust gate and the “letter absent” outcome deserve their own checks because a happy-path visit would not verify them. Each row should point to the relevant node or transition in the graph so a changed condition can be traced to affected tests.

Section 4

Prioritize branch coverage when combinations grow

If a story has many independent flags, the number of possible combinations can grow quickly. Do not respond by listing every theoretical route as a mandatory full playthrough. First identify high-risk edges: choices that gate endings, consume items, set persistent relationship facts, or merge histories. Test those transitions and their important downstream outcomes directly.

Then sample combinations deliberately. Include boundary conditions (the minimum value that changes a choice), both sides of each critical condition, representative combinations of flags that can interact, and routes that reach a rejoin point through different histories. Prioritize recent changes and paths with complex prerequisites. When two variables may interact, add a test for that pair rather than assuming separate single-variable checks prove the combination works.

Record the coverage unit and its limits. A team might track whether every critical choice edge was exercised, whether each condition was checked both true and false where relevant, and whether every ending trigger was reached by at least one designed test. Those are useful reports of what was sampled; none proves every possible history, state combination, or wording issue has been explored.

Research on game playtesting can offer a related but bounded idea. Gordillo and colleagues’ 2021 arXiv paper describes reinforcement-learning agents rewarded for novel actions to explore state coverage in a complex 3D scenario. That work concerns exploration in a 3D game environment, not branching narrative choices or the specific transition-test method described here. It supports treating automated exploration as a possible complement, but neither that study nor Tikhonov’s paper validates this exact narrative test method. [Gordillo et al., “Improving Playtesting Coverage via Curiosity Driven Reinforcement Learning Agents”](https://arxiv.org/abs/2103.13798)

Section 5

Add negative assertions and human playtests

Positive assertions confirm that the expected choice or outcome exists. Negative assertions confirm that something forbidden does not happen: a locked option does not appear, a consumed clue is not granted twice, an absent letter does not trigger its line, or a failed attempt does not set a success flag. Negative checks are especially helpful around shared nodes, where stale state from another route can leak into the current scene.

Automated checks can verify route logic and exact state changes, but they do not reliably judge whether a transition feels coherent, whether a line contradicts what the player remembers, or whether a choice is understandable in context. Human playtests should therefore use selected routes with a purpose: ask testers to follow a less common branch, arrive at a rejoin with a particular history, or try to reach an ending while missing a key. Observe both the resulting state and the player’s interpretation of it.

Keep playtest notes tied to node and choice identifiers, plus the initial state and steps taken. That makes a reported issue reproducible and helps distinguish a writing concern from a logic defect. After a fix, rerun the affected transition tests and at least one representative route through the changed rejoin or outcome.

Section 6

A practical test cycle for a changing story

For each story update, export or review the graph, identify changed nodes, conditions, effects, and rejoin points, then update the coverage matrix. Run focused state-transition tests first; follow with selected route samples and human playtests for high-impact or newly changed sections. When a failure appears, capture the starting state and choice sequence so the team can reproduce it, repair the relevant rule or passage, and preserve the case as a regression check.

The goal is a testable account of what has been checked and why. A graph makes structure inspectable, transition tests make logic explicit, branch sampling directs effort toward meaningful variation, negative assertions catch leaked state, and human playtests assess interpretation. Together they provide useful coverage as paths multiply, while leaving a clear boundary around what remains untested.

Related reading

Keep exploring this topic