• AI-assisted workflows

What are good human-in-the-loop UX patterns?

11 Min Read11 Min Read

Last updated on 5 Oct ‘26

Insights

Eight human-in-the-loop UX patterns for AI products, from evidence-first approval to undo windows and override records, with the research behind each. Good human-in-the-loop patterns give the reviewer evidence next to the proposed action, a real choice to approve, edit or reject it, and a way back when it is wrong.

Good human-in-the-loop patterns give the reviewer evidence next to the proposed action, a real choice to approve, edit or reject it, and a way back when it is wrong. Match the pattern to the consequence: a bare approve button invites rubber-stamping, while review on every low-risk action wears people out.

Last reviewed October 5, 2026.

This page is a pattern library for the review step in an AI product. It covers eight patterns, ordered by when they act: before the system commits, while it runs, and after it has acted. To see how nine shipping products place the human check, read AI trust, control and human review: nine examples.

Why is an approve button not enough?

Putting a person in the loop does not by itself produce oversight. Four findings explain why, and each points at something a designer can change.

People over-rely on automated output. A review of the human-factors literature found that automation bias appears in both naive and expert participants and is not removed by training or instructions (Parasuraman and Manzey, 2010). The practical reading is that reminding reviewers to be careful is a weak control. The interface has to carry the load.

Explanations can raise acceptance of wrong answers. In user studies where the AI was about as accurate as the people using it, adding explanations did not improve combined performance. It increased how often people accepted the AI's recommendation, whether or not it was right (Bansal et al., CHI 2021). A later set of five studies (731 participants) found a boundary condition: explanations reduced over-reliance when verifying the AI's answer was easy enough to be worth the effort (Vasconcelos et al., 2023). The design lesson is to lower the cost of verification, not to add more justification text.

Repeated approvals wear attention down. Anthropic reports that Claude Code users approve 93% of permission prompts and describes the result as approval fatigue, where people stop paying close attention (Anthropic Engineering, March 2026). Two cautions apply. A high approval rate is also what you would see if most requests were fine, so the number alone does not prove careless review. And this is a coding agent's permission prompts, which is neighboring evidence for a B2B product's review queue, not direct evidence about it.

Policy assumes a capability the evidence may not support. A survey of 41 government policies that require human oversight of algorithms argues that people often cannot perform the oversight functions assigned to them, and that the requirement can create a false sense of security (Green, 2022). That is an argument about government algorithms, and it is contested territory, but it sets the right bar for product teams: prove the review role is performable.

Regulation points the same way. For high-risk systems, Article 14 of the EU AI Act requires that the people assigned to oversight can understand the system's limits, stay aware of automation bias, interpret the output, decide not to use it or to override or reverse it, and stop the system in a safe state (EU AI Act Service Desk, Article 14). Whether or not the Act applies to your product, that list is a useful design brief.

What are the eight patterns?

Before the system commits

1. Evidence-beside-proposal approval

Use when the action is consequential or hard to reverse.

What the reviewer sees: the proposed change, the source it rests on, the objects it affects, the current state, and the consequence of approving, with the reason for review shown before the controls. The Model Context Protocol specification, an open standard for connecting AI applications to external tools, says clients should show tool inputs to the user before a call and applications should insert clear visual indicators when tools are invoked (MCP specification).

Design judgment: show the source material itself, with the relevant passage marked, rather than the model's account of why it is right. That follows from the Bansal and Vasconcelos findings above, but it is our inference, not a tested result.

Add a freshness condition where inputs can change. An approval covers a bounded decision: the evidence inspected, the affected object and the state at that moment. If price, inventory, policy or the record itself can change before execution, the system should revalidate and send stale work back to the reviewer instead of executing it.

2. Approve, edit or reject as equal outcomes

Use when the reviewer can fix a near-miss faster than they can send it back.

Orchestration frameworks already model this. LangChain's human-in-the-loop documentation lists three decisions on a proposed tool call: approve it as is, edit its arguments, or reject it with feedback (LangChain human-in-the-loop). It builds on LangGraph interrupts, which pause a run and save its state (LangGraph interrupts). The UX work is making all three outcomes equally easy. If editing takes far more effort than accepting, the interface quietly rewards weak approval. And rejection needs somewhere to go: an alternative, a manual path, or a return with a recorded reason.

3. Make the reviewer commit first

Use when the decision is high-stakes and wrong approvals are costly.

Buçinca, Malaya and Gajos tested three "cognitive forcing" designs with 199 participants: showing the AI suggestion only on request, having people decide before seeing it, and delaying it by 30 seconds. Compared with simple explainable-AI baselines, forcing reduced over-reliance. But people gave the designs that cut over-reliance the most the least favorable ratings, and the benefit was larger for people higher in Need for Cognition (Buçinca et al., 2021). The task was a controlled online experiment, not a professional workflow. Reserve this friction for the decisions that justify it, and measure whether your reviewers tolerate it.

4. Route by risk and exception

Use when most output is routine and a few items are not.

Sending every result to one queue can bury the exceptions. Route on checkable conditions: risk, missing sources, policy or permission conflicts, and repeated disagreement. Show uncertainty where it changes the decision, but do not expect a confidence score to do the reviewer's work: in two experiments, confidence scores helped calibrate trust, yet calibrated trust alone did not improve decisions (Zhang, Liao and Bellamy, 2020).

Also remove repeats. In a study of 112 primary-care clinicians, acceptance of reminders fell by about 30% for each additional reminder per encounter, and by about 10% for each five-point rise in the share of repeated reminders. The authors found no sign that a newly added alert wore out over time (Ancker et al., 2017). It is clinical software, so treat it as neighboring evidence, but it supports de-duplicating and batching the asks before adding more of them.

5. Pre-approve a scope, review the exceptions

Use when actions are low-risk, reversible and inside a permitted boundary.

Instead of a prompt per action, the person approves a scope once, and the system interrupts only outside it. Anthropic's auto mode is one example: a classifier decides which tool calls need a human. Anthropic reports a 0.4% false-positive rate on 10,000 real tool calls from Anthropic employees and a 17% miss rate on a curated set of 52 real over-eager actions, and states that it is not a drop-in replacement for careful human review on high-stakes infrastructure (Anthropic Engineering). Those are one vendor's measurements of its own product. The transferable point is that an automated gate will miss some actions, so keep the pre-approved scope to things that can be corrected, and keep a trace.

While the system runs

6. Stop and takeover

Use when the system runs several steps, or runs without a person watching.

Article 14 asks for a stop control that brings the system to a halt in a safe state. In product terms, the stop must not leave half-finished work. Where one step succeeds and the next fails, show what changed, keep the valid progress, and check the external result before offering a retry that could duplicate an action. A revised plan may need a fresh approval. Pausing also has a technical dependency: the paused state has to be stored durably, which is why LangGraph requires a persistence layer for interrupts.

After the system has acted

7. Undo window for reversible actions

Use when the action can be deferred or reversed, and confirming each one would train people to click through.

Aza Raskin's argument is that confirmation dialogs fail because people form habits and click through them, so "never use a warning when you mean undo" (Raskin, A List Apart, 2007). Gmail's undo send is the familiar case: a short delay, 5 seconds by default and up to 30, before the message leaves (Google). The mechanism matters. Undo works because the side effect is held back or can be reversed. If the effect reaches another system or person immediately, an undo button is a promise the product cannot keep, and the review has to move before the action.

8. Override with a reason, and keep the original

Use when reviewers will disagree and the disagreement is useful.

A correction should record the new state and a reason, and the system's original result should stay visible next to it. Silent replacement makes the interface look simpler and erases what the approver, the product team and later evaluation need. In the BuildTwin case, Tcules designed review of AI quality checks this way: results linked to their source evidence, missing-input states, and overrides that carried a reason and stayed in the record. That case does not establish production workload, adoption or operating results, so read it as a design pattern, not a measured outcome. Microsoft's guidelines for human-AI interaction include the same ideas as support for efficient correction and granular feedback (Amershi et al., CHI 2019). When the same disagreement keeps recurring, send it to model or product evaluation instead of leaving it as a pile of overrides.

Which pattern fits which decision?

SituationLead patternSupport with
High consequence, hard to reverseEvidence-beside-proposal approval (1) with approve, edit or reject (2)Commit-first friction (3) where a wrong approval is costly
Material uncertainty or a missing sourcePause and expose the evidence gap (1)Routing to the person who can supply context (4)
Low risk, reversible, inside permitted scopePre-approved scope (5)Undo window (7) and a trace (8)
Policy or permission conflictBlock the action and explain why (4)A route to an authorized person
Multi-step run goes wrong partwayStop and takeover (6)Preserved partial state, reconcile before retry
Repeated disagreementOverride record (8)Escalation to product and evaluation review

The rows are a starting point, not a rule. In our judgment, reversibility and who can step in should weigh at least as much as the model's confidence.

How do you test that review actually works?

An approval rate is not a quality measure. Anthropic's 93% would look the same whether reviewers were careful or tired. Test the review step directly:

  • Catch rate. Seed the queue with known-bad items in a test and see how many reviewers stop. This is our suggested method, not one taken from the studies above.
  • Disagreement rate and reasons. Zero disagreement is, in our view, a warning sign, not a success.
  • Attention per item. Time and effort for approve versus edit. A big gap says the interface favors acceptance.
  • Absent reviewer. What happens at 2 a.m. or on leave: does work wait, escalate to someone with authority, or continue without the AI path?
  • Stale approvals. Do approvals expire when the underlying state changes?
  • Repeat asks. How many reviewer prompts are duplicates?

If you can answer what the person is checking, what evidence they receive, what they can change, what happens on rejection and what responsibility transfers on approval, the loop has a defined job. If you cannot, the person in it is mostly there for assurance.

Where does this work sit?

Choosing among these patterns depends on what your product's decisions cost when they go wrong, which is a design question before it is a model question. Tcules designs this step as part of its Human-in-the-Loop Workflows service: the review job, the evidence, the correction path and the recovery path. If you want to see how other products place the human check, AI trust, control and human review: nine examples walks through nine of them.

Tell us about the product problem you are working on.

Talk to Tcules fast and affordable

Start a project