Metlivi Blog

How to Measure a Useful Short AI Chat Without Optimizing for Session Length

A short AI conversation can be successful when it helps someone finish a small creative task: choose a color palette, name a weekend project, or find a few ideas for a personal activity. To evaluate that exchange, define what the person wanted, then check whether the response was relevant, useful, and easy to steer or end. Time spent and return visits can describe usage, but by themselves they cannot tell you whether the exchange helped.

September 30, 20266 min readReading, Arts & CultureBy Metlivi Editorial Team
Section 1

Start with the task the person came to do

Before choosing a metric, name one ordinary task in observable terms. “Get a few playful names for a balcony herb garden” is easier to assess than “have an engaging conversation.” The first gives reviewers something concrete to look for: did the system offer names that fit the request, and could the person select or adapt one?

This follows a broader usability principle: judge whether specified users achieve specified goals effectively, efficiently, and with satisfaction in a particular context. The ISO definition treats these as related but distinct qualities, so no single measure—duration included—stands in for all of them. ISO 9241-11:2018, *Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts*

For creative requests, define completion as a useful outcome rather than a single correct answer. A person might leave with one chosen idea, a shortlist, or a better direction for another attempt. Write down in advance what counts as enough for the task. That prevents teams from quietly redefining success as “the user kept chatting.”

Section 2

Score completion and relevance separately

A practical review can start with two questions: did the exchange reach the stated outcome, and did the assistant’s replies fit the request? Keep the answers separate. A response can be relevant but leave the person without a usable choice; a person can also complete a task after ignoring a distracting or generic reply.

Task completion is a common, simple measure, but it compresses outcomes into a yes-or-no result and does not explain why someone failed or how well a completed task went. Pair it with a short review of the conversation or a direct user rating. Nielsen Norman Group, “Success Rate: The Simplest Usability Metric”

Dialogue research supports this more detailed view. A study of task-oriented dialogue annotated relevance, interestingness, understanding, task completion, and efficiency at both turn and conversation levels; it found that a single overall rating can miss useful distinctions, and that satisfaction varied across annotators and dialogues. Those findings are not a universal scoring formula, but they offer a grounded set of dimensions to adapt to a creative task. Siro, Aliannejadi, and de Rijke, “Understanding User Satisfaction with Task-oriented Dialogue Systems”

Section 3

Treat response quality as a turn-level question

Review whether each answer responds to the latest request and uses relevant context from earlier turns. If the person says “make those less formal,” for example, a useful next reply should change the suggestions accordingly. Count avoidable repetition, missed constraints, or requests for clarification that do not help move the task forward.

A study of intelligent assistants found that satisfaction differed by scenario: task completion mattered strongly in some tasks, while effort mattered in others. It also reported that preserving conversational context was important and that overall task satisfaction could not be reduced to satisfaction with individual queries. The practical implication is to inspect both the particular replies and the whole exchange, while interpreting each against its task. Microsoft Research, “Understanding User Satisfaction with Intelligent Assistants”

For a creative exchange, a short human review might label replies as relevant, partly relevant, or off task, and note whether the assistant followed corrections. These labels are a proposed review method, not a validated benchmark. If an automated classifier supplies them, check samples manually: a tidy score can still miss a suggestion that technically matches the prompt but gives the person nothing they can use.

Section 4

Check that the person could steer the exchange

User control is visible in small moments: the person can narrow the request, ask for another style, correct a misunderstanding, or stop after getting enough. Review transcripts for whether the system responded to direction instead of repeatedly pursuing its own line of conversation. If the interface provides controls to stop, edit, regenerate, or clear a response, test that those actions behave as people expect.

Research on conversational agents has described autonomy in terms that include control over the conversation, questions, and responses. This supports treating control as a quality to inspect, though it does not establish one universal product metric for it. Yang and Aurisicchio, “Designing Conversational Agents: A Self-Determination Theory Approach”

A lightweight user question can make control measurable: “Could you get the conversation to the kind of result you wanted?” Ask it after the exchange, alongside a task-completion question. Keep the response options consistent across sessions and allow a short explanation. A low score is a signal to inspect the conversation, not proof of a particular cause.

Section 5

Recognize a clear ending as a valid outcome

A good endpoint is when the person has what they need and can leave without another prompt from the system. For the herb-garden example, that might be the moment they choose a name. A conversation that ends there may be more useful than one extended by generic follow-up questions. This is a measurement principle inferred from the task: if the goal is complete, extra turns do not automatically add value.

Track whether the task appears complete and whether the person chose to continue, but interpret those events in context. A follow-up turn could reflect curiosity, a necessary correction, or an incomplete answer. A silent exit could mean the person is satisfied, distracted, or disappointed. Conversation length alone cannot distinguish these explanations; use sampled transcript reviews or brief user feedback to investigate.

Section 6

Use duration as context, not the verdict

Duration can help answer operational questions, such as whether a particular flow routinely requires many turns. It is still a measure of elapsed activity, not direct evidence that an answer was useful. Google Analytics, for example, defines engagement time as the duration of user engagement associated with events; its developer guidance also warns that artificially extending sessions distorts behavior data. These are analytics definitions and implementation cautions, not a quality measure for creative chat. Google Analytics, “Measurement Protocol reference” and Google Analytics, “Measure sessions and user engagement”

Compare time only among similar tasks and alongside task completion, relevance, user control, and a clear endpoint. A shorter median duration might reflect a more direct answer, but it could also reflect users giving up. A longer duration might mean productive iteration—or unnecessary friction. The supporting evidence has to come from the task outcome and what people experienced, not the clock alone.

Section 7

A compact evaluation routine

For a small study, give participants one clearly worded creative task, let them use the chat naturally, and record whether they reached the outcome they defined. Review a sample of exchanges for relevance, context-following, and opportunities to steer or stop. After each task, ask whether they got a useful result and felt able to direct the conversation. Record duration as context, then inspect cases where duration and reported usefulness disagree.

Report the measures separately instead of collapsing them into one “engagement” score. State the task, how completion was judged, who reviewed replies, and how user feedback was collected. If the sample is small or the task definition is subjective, describe the findings as directional rather than generalizing them to every user or creative request.

A short exchange matters when it carries the person from a specific request to a result they can use, while leaving them in control of the route and the stopping point. Measure those outcomes directly. Let session length help explain the experience, but do not let it define success.

Related reading

Keep exploring this topic