Almost every AI automation proposal says a human stays in the loop. Far fewer say what the human is actually supposed to do, or design the workflow so they can do it.
An approval button that everyone clicks without reading is not oversight. It is a liability with a nice interface — it creates a record suggesting a person checked, when in practice nobody did.
Here is what we have found actually works.
Put the approval where the risk is, not everywhere
The instinct after an incident is to add review steps. The result is a workflow where a person confirms twelve things, eleven of which never vary.
That trains people to click. By the time the twelfth thing — the one that matters — arrives, the habit is already automatic.
Approve fewer things, more carefully. Let the routine cases run. Route the unusual ones to a person with a clear explanation of why they were flagged.
Make the exception explain itself
“Requires approval” tells someone nothing. It gives them no way to decide quickly, so they either rubber-stamp it or spend ten minutes reconstructing the context by hand.
Compare that with: “This quote is 40% above the usual range for this route. The vehicle is listed as non-standard dimensions.”
Now the reviewer knows what to look at. The review takes twenty seconds and is a real review.
The rule we work to: if the reviewer has to open another system to make the decision, the workflow has failed them. Put what they need in front of them.
Show the working, not just the answer
When AI drafts something — feedback, a reply, a summary, a classification — the person reviewing it needs to know what it was based on.
For document extraction, that means showing the field alongside the part of the document it came from. For a drafted reply, it means showing the source information. For a classification, it means showing which signals drove it.
This is not a transparency nicety. It is the difference between someone being able to spot a wrong answer in five seconds and having to redo the work to check.
Make disagreeing easy, and record it
If correcting the AI is harder than doing the task from scratch, people will quietly stop using the system, and you will not find out for months.
Editing a draft should be a text box, not a rejection workflow. And every correction is a signal — a step producing corrections 30% of the time is telling you something about either the automation or the process behind it.
Track that. Rising correction rates are usually the earliest warning that something upstream has changed.
Keep an audit trail people can actually read
“An audit trail” often means logs a developer can query. That is not much use during an audit, a complaint, or a conversation with a regulator.
What you want is: for any given case, in plain language, what happened, when, which steps ran automatically, what the AI produced, who approved it, and what they changed. Available to the people who need it, without a database query.
This is far easier to build in from the beginning than to retrofit.
Decide what happens when the automation is unsure
Every automated step will eventually meet a case it cannot handle. The question is what it does then.
The wrong answer is to guess and continue. The right answer is usually to stop, hold the case, and tell a person — with enough context for them to act.
That means designing the “I do not know” path as carefully as the happy path. In practice this is where most of the engineering effort goes, and it is what separates automation that survives contact with real work from a demo.
Why this matters commercially, not just ethically
Automation that people do not trust gets worked around. Someone keeps a parallel spreadsheet. Someone checks every case manually “just to be safe”. The efficiency you paid for quietly evaporates while the system still reports that it is running.
Designing genuine oversight is not a tax on the automation. It is the thing that lets the automation be used at full strength, because people know where its limits are and trust it inside them.
If you are thinking about where the approval points should sit in a process of your own, that is exactly what we map first.