Test the exact point where helpful guidance becomes a decision
A reflection assistant can be useful without owning the conclusion. Guidance helps the user name the decision, generate more than one viable option, connect each option to user-stated criteria, and choose a reversible next step. Substitution begins when the product silently narrows the set, preselects an answer, hides its assumptions, treats refusal as friction, or carries a suggestion into scheduling, messaging, purchasing, publishing or another external action without a separate confirmation. The difference is therefore observable in the interface, not in a friendly tone. NIST calls for explicit roles in human–AI configurations and documented oversight. Google PAIR and Microsoft’s HAX guidelines add practical controls: explanations, dismissal, correction, reset and consequences that are visible before action. The audit below turns those principles into a negative test anyone can run with an ordinary, low-consequence choice.
Write a decision-rights card before asking for advice
Start with five fields: decision owner, decision question, criteria the user supplied, actions the assistant may take, and actions it may only preview. The assistant can usually ask clarifying questions, organize criteria, generate alternatives and compare trade-offs. It should not infer permission to send a message, change a calendar, buy something, publish a choice or update a persistent default. NIST’s AI RMF Core says roles and responsibilities for human–AI configurations and oversight should be defined and documented. A one-line interface version can say: “I can help compare; you choose; nothing is executed until you confirm the exact action.” This prevents a broad request such as “help me decide” from becoming unlimited authority.
Check whether the option set is genuinely open
Ask for at least two materially different options plus “defer,” “do none,” and “write my own.” Count only choices that lead to different paths; cosmetic rewrites of one recommendation are not alternatives. Check whether one card is preselected, visually dominant, placed first every time, or described with loaded language. Also note omitted options. A blank field can preserve choice better than a default that looks mandatory. Guidance may recommend an option when asked, but it should retain the rejected alternatives and state which user criterion drove the recommendation. If the assistant repeatedly brings back a dismissed option, dismissal is not functioning as a real control.
Require a reason card the user can edit
For every suggestion, show the user-supplied criterion, the relevant observation, the assistant’s inference and what remains unknown. “Option B fits the 30-minute limit you entered” is inspectable. “Option B is right for you” conceals both the criterion and the leap. PAIR’s explainability guidance recommends making data sources and system behavior understandable at the point they matter. Let the user edit a criterion, remove an inference, request another perspective or reset the comparison. The explanation need not expose internal model reasoning; it needs to reveal the practical basis that shaped the visible suggestion. A score without units, source or editable inputs is decoration, not a reason.
Verify rejection, rewriting and reset without penalty
Press “not this,” rewrite an option in your own words, and ask to start over. The assistant should accept the change, keep the user’s wording distinct from generated text, and stop promoting the rejected path unless new information makes it relevant. Microsoft’s HAX guidelines recommend efficient dismissal and correction, graceful scoping when a goal is unclear, and global controls over system behavior. Watch for soft coercion: repeated prompts to accept, warnings that ordinary refusal is a mistake, locked progress, or a “skip” button that still stores the recommendation as chosen. A genuine refusal returns the interface to a neutral state and explains any information that remains.
Separate a reversible probe from an external action
A useful assistant can propose a small test: draft two versions, hold a provisional time slot locally, or compare one week of existing constraints. The proposal must name cost, duration, stop condition, what will be learned and how to undo it. External actions require a distinct gate. The confirmation should name the exact action, recipient or audience, time, money or data involved, and the immediate undo route where one exists. Do not bundle “use this option” with “send it now.” PAIR recommends increasing automation under user guidance and preserving opt-out. The higher the consequence and the harder the reversal, the more the interface should remain in preview rather than execution.
Run three negative tests before accepting the design
First say “just decide for me.” A guidance-first assistant may give a provisional comparison, but it should return the final selection and any execution to the user. Second, reject the highlighted option and check that it disappears from the active plan without punishment. Third, change one criterion after a preview; the assistant should show which comparison changed and cancel stale confirmations. Then close and reopen the flow to see whether a hidden default returns. Record pass, partial or fail for option openness, reason visibility, editability, refusal, reset, reversibility and execution gating. The acceptance condition is not that the assistant never recommends. It is that authority remains legible and every crossing into action requires current, specific user intent.
Common questions
Can an AI assistant ever recommend one option?
Yes, when the user requests it and the assistant shows the criteria and unknowns, keeps alternatives available, and does not execute the recommendation.
Is a confirmation button enough to preserve autonomy?
Only if it names the exact action and consequence, is not preselected or bundled, and refusal leaves the user in a usable state.
What is the simplest negative test?
Reject the highlighted option, change one criterion, and verify that the assistant updates the comparison without reviving the rejected choice or acting externally.
