Digital employee

Supervising AI agents: your team's new job and how to do it well

A report published this week confirms what we see in every project: when an AI agent arrives, people's work shifts from doing to supervising. But supervising isn't approving whatever appears on screen. Here's how to design oversight that actually controls.

Amud team · Reviewed on October 9, 2026 · 4 min read

What the report says

The report Agentic AI: the new human-machine alliance by MIOTI Tech & Business School, published this week and based on more than 70 technology and HR leaders across 25 sectors, includes several findings reported by RRHH Digital:

  • 7 in 10 leaders want people to stop executing tasks and move to supervising and controlling agents.
  • 37% see the person as a strategic supervisor evaluating the final result; 33% as an operational gatekeeper authorising each critical step.
  • 73% are still worried about model "hallucinations", and 90% want agents that remember the working context.

On the customer side the message is similar: according to an Experian study published in September, 69% of Spaniards want transparency and human oversight before trusting AI agents.

The conclusion is clear: the agent does the work and the person answers for it. The question is how to organise that oversight so it's real and not a formality.

Two ways to supervise

The report's two positions aren't mutually exclusive. They're two tools, and a good design uses each where it fits.

Operational gatekeeper: approve before it happens

The agent prepares the action and stops. A person reviews it and decides whether it runs. This fits when the action:

  • Affects people: a dismissal letter, turning down an application, a claim against a customer.
  • Is hard to undo: a sent email, a payment, a filed document.
  • Commits the business: a quote, a commercial condition, a reply with legal implications.

Strategic supervisor: review the result

The agent acts on its own and the person reviews afterwards, by sampling and with metrics. This fits low-risk, high-volume, easy-to-correct work: sorting email, requesting documents, updating data, answering FAQs.

How to decide what goes where

For each action the agent can take, ask three questions:

  1. What happens if it gets it wrong? If the harm is serious or affects a person, prior approval.
  2. Can it be undone? If not, prior approval.
  3. How many times a day does it happen? If hundreds and the risk is low, sampling.

The result is a simple table, action by action, that becomes part of the agent's job description. We cover it in how to implement a digital employee.

The big risk: rubber-stamping

When someone approves a hundred proposals a day and almost all are fine, they end up approving out of habit. It's called automation bias, and the EU AI Act itself mentions it when regulating human oversight of high-risk systems. These measures reduce it:

  • Approvals with context. Each proposal shows what the agent will do, why and with which data, so a decision takes seconds without opening five programs.
  • Fewer, better-chosen approvals. If everything needs approval, nothing really gets reviewed. Low-risk work moves to sampling.
  • Control cases. Every so often, review a random proposal in depth, even if it looks right.
  • Single execution. What's approved runs once and exactly as approved; if anything changes, it asks again.
  • Rotation and limits. The same person shouldn't approve for hours on end.

What to measure

Oversight is managed with data, not gut feeling:

MetricWhat it tells you
Share of proposals corrected or rejectedIf very high, the agent needs adjusting; if zero for months, people may be rubber-stamping
Time to approvalIf it grows, oversight has become a bottleneck
Errors found in samplingThe real quality of what the agent does alone
Cases escalated by the agent itselfWhether it recognises what isn't its job
Hours recoveredThe project's return

With these figures you decide when to give a task more autonomy and when to take it back.

The person's role changes too

Supervising agents is skilled work. The person who does it well is usually whoever used to do the task, because they know what a good result looks like and spot what's off. They need:

  • To know the agent's job description: what it does, what it doesn't and what it must check.
  • To know how to correct: proposing changes to instructions and rules, not just rejecting.
  • Specific training, which is also part of the AI literacy the AI Act requires (we explain it in the EU AI Act after the Omnibus).

How we do it at Amud

In the digital employees we build, oversight isn't bolted on at the end: it's part of the design. Each request starts with a written success criterion; anything with consequences stops and waits for approval; what's approved runs only once; and everything is recorded in the case file with its author. That way the team moves from doing the work to directing it, without losing control.

Sources

Frequently asked questions

What does supervising an AI agent mean?

A person decides what the agent can do on its own, approves actions with consequences before they run, reviews a sample of what it does without approval and corrects its instructions when something goes wrong. It's a job of judgement, not constant watching.

Do you have to approve everything an agent does?

No. Approving everything creates bottlenecks and, over time, approvals without looking. The usual approach is to always approve what affects people or is hard to undo, and to sample-check low-risk, high-volume work.

How do you avoid rubber-stamping?

By presenting each approval with the information needed to decide (what the agent will do, why and with which data), limiting the volume each person receives, adding occasional control cases and measuring how many proposals get corrected or rejected.

Does supervising agents require technical knowledge?

No. It requires knowing the process well: knowing what a good result looks like and spotting when something's off. That's why the best supervisor is usually whoever used to do the work.

Keep reading

Which task would you like off your plate?

Tell us how your team works. In a free 30-minute session we'll tell you what can be automated, how much you'd save and what isn't worth it.

Book a free meeting