Metlivi Blog

How Game Developers Can Estimate the Cost of Long AI Dialogues

To estimate what long AI conversations will cost in a game, measure the tokens used across complete player sessions, separate uncached input, cached input, and output, then apply the selected model’s current rates. Finally, weight the result by how many sessions players actually have at each length. A short prompt test or a single “average” session can miss the repeated history sent on later turns, cache misses, and unusually long play sessions.

September 30, 20266 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

1. Define the unit you are estimating

Choose a clear unit first: for example, model inference cost per dialogue session, per active player-day, or per 1,000 sessions. For a session estimate, define when the session starts and ends. A practical rule might be “from the first dialogue request until 30 minutes without a request,” but that timeout is a measurement choice, not a universal standard. Record it so another developer can reproduce the estimate.

Count model requests, not just player messages. A single interaction can trigger several calls—for example, a response followed by a separate tool-handling call—and retries can add more. If the game uses voice, image input, retrieval, or tools, keep those in separate cost fields as well as recording their associated model tokens. Provider pricing can include tool fees or modality-specific rates in addition to ordinary text token charges; OpenAI, for example, lists separate tool charges and says model tokens used for built-in tools are billed at the chosen model’s rates (OpenAI API pricing).

Section 2

2. Measure actual tokens per request

For every call, log the provider, model identifier, timestamp, session identifier, request purpose, input-token count, output-token count, and any reported cached-token or reasoning-token breakdown. Capture retries, errors, tool calls, and whether the call completed. Avoid storing dialogue content unless it is needed for a separately justified product purpose; token totals and operational metadata are usually enough for a cost estimate.

Use provider-reported usage from completed calls as the primary measurement. Preflight counting is helpful for testing prompt construction, but it may not match final billing fields. OpenAI documents that reported output includes tokens beyond visible text, such as some formatting and tool-structure tokens, and recommends not estimating output solely from what a player sees (OpenAI token-counting guide). Google’s Gemini documentation likewise distinguishes prompt, cached-content, candidate-output, and thinking token counts in usage metadata (Gemini token guide).

Build a representative sample that includes new players, returning players, short and long conversations, and the production prompt and tool configuration. Keep sessions as the sampling unit: a 40-turn dialogue should remain one observation with its accumulated cost, rather than being treated as 40 independent player sessions. During a prototype, a fixed set of scripted conversations helps compare prompt changes; after launch, observed sessions should drive the forecast.

Section 3

3. Account for history growth and context reuse

In many dialogue systems, each request includes the current user turn plus some or all of the conversation history. If history is repeatedly resent, input tokens can grow with each turn even when each player message is brief. Measure the actual request payload sent to the model; do not multiply one-turn prompt size by the turn count unless the implementation really sends the same amount each time.

Caching changes the rate applied to eligible repeated input; it does not mean the full dialogue becomes free or that a persistent session guarantees a cache hit. OpenAI describes prompt caching as reuse of an unchanged prompt prefix and notes that new input still has to be processed. Its cache diagnostics can help measure cache reads and misses (OpenAI prompt-caching guide). Anthropic similarly distinguishes cache writes from cache reads and publishes separate rates and cache durations (Anthropic pricing and prompt caching).

In your telemetry, separate uncached input tokens, cached input tokens, and cache-write tokens when the provider reports them. Record cache hits divided by eligible requests, and the share of input tokens actually billed as cached. These answer different questions: a high hit rate across requests can still mean a modest share of total tokens was cached if the repeated prefix is small. Cache eligibility, minimum sizes, expiry, prompt-prefix stability, and model support are provider-specific; count only savings confirmed by usage data.

Section 4

4. Apply rates with a transparent calculation

For a model priced per million tokens, calculate each session as:

session cost = (uncached input × input rate + cached input × cached-input rate + cache writes × cache-write rate + output × output rate) ÷ 1,000,000 + other applicable charges

Use the rates for the exact model, endpoint, modality, region, and service tier in the deployed configuration. Check them independently on the provider’s pricing page immediately before preparing a budget; rates and model catalogs change. Include storage charges where a cache bills for stored tokens and duration. For example, Gemini’s published pricing lists token categories for paid use and storage-hour pricing for certain context-caching configurations, while its billing documentation identifies input, output, cached tokens, and cache-storage duration as billable factors (Gemini pricing; Gemini billing).

A worked calculation makes assumptions visible. Suppose, purely for illustration, that measured calls assigned to one session contain 18,000 uncached input tokens, 12,000 cached input tokens, and 6,000 output tokens. Applying Gemini 3.8 Flash paid-tier rates published for use through December 31, 2026—$0.75 per million input tokens, $0.075 per million cached tokens, and $3.75 per million output tokens—gives $0.0135 + $0.0009 + $0.0225, or $0.0369 before any applicable cache storage or other charges. This example assumes the listed cached tokens are billed at the cached rate and does not include any initial cache creation cost outside the measured totals. The rates are time-bounded and should be checked again when estimating a later period (Gemini pricing).

Section 5

5. Use the distribution of session lengths

Do not multiply one hand-picked “typical” session by total players and call it a forecast. Group observed sessions by turn count or another useful length band, calculate the average cost within each band, then weight each band by its share of sessions. Keep the median and upper percentiles alongside the mean: the mean estimates total usage when multiplied by session volume, while percentiles help describe what a shorter or unusually long session can cost.

For example, if a sample has many short sessions and a small number of very long ones, report both the share of sessions in each band and each band’s cost. A forecast for 10,000 sessions can then be calculated as the sum of sessions in band × mean cost in band, rather than assuming every session resembles the overall median. If usage differs materially by game mode, language, platform, or new-versus-returning player, stratify those groups before combining them. The choice to segment is an analysis decision; explain why each segment could change token usage or request behavior.

Section 6

6. Document uncertainty and refresh the estimate

Keep a compact assumptions record with the sample dates, session boundary rule, models and endpoints, prompt version, observed session counts, cache-hit definition, price-page retrieval date, included fees, and excluded components. Show a low, central, and high scenario by changing observable inputs—for example, session length mix, output size, or measured cache-hit share—rather than applying an unexplained buffer. Treat any projected behavior beyond observed sessions as an explicit scenario, not a measured fact.

The estimate is only as complete as its instrumentation and billable categories. It can miss calls routed outside the logger, retries, cached-token categories that the API does not expose, model-side reasoning tokens, storage duration, media processing, or provider charges beyond token rates. Reconcile sampled usage against provider billing reports when available, investigate material gaps, and rerun the calculation after changing models, prompts, caching behavior, or player-facing features. The result is a documented operating-cost estimate for the observed workload, not a promise that future sessions or invoices will match it.

Related reading

Keep exploring this topic