How much should you check the AI’s work?
Author
Richard S O'Rourke
Date Published
Two things go wrong when a practice starts using AI in earnest, and they look like opposites. One is work that arrives looking finished and isn’t — someone downstream has to redo it. The other is the person doing the checking wearing out. Harvard Business Review has now put numbers on both. The instrument below is how I read them together.
Two workplace phenomena documented in Harvard Business Review:
Workslop
AI output that looks polished but lacks the substance to advance the task, offloading the real work onto whoever receives it.
A survey of 1,150 full-time U.S. desk workers found 41% had received workslop in the past month; each instance took an average of 1 hour 56 minutes to resolve — an estimated $186 per employee per month, and over $9M a year for a 10,000-person organisation.
A follow-up locates the cause less in individuals than in management: mandating AI use without training or clear guidelines produces the low-effort output, and stricter reviews or harsher feedback don’t reduce it.
Brain fry
Acute mental fatigue from marshalling cognitive oversight of AI beyond capacity — the cost of judging it all by hand.
A survey of 1,488 full-time U.S. workers found 14% of AI users had experienced brain fry. Affected workers made 39% more major errors and 11% more minor errors, showed 33% more decision fatigue, and their intent to quit rose from 25% to 34%.
The evaluation performance loop · how good the output gets vs how hard you check — your skill level is the current ceiling (it rises as you learn); the amber line is what the stakes demand
● near the sweet spot — sustainable
your skill 50/100 · Competent · at Competent you can stand behind output up to 50% of what a perfect judge would catch — above what these stakes demand, so you can sign off on the verdict. That ceiling isn’t fixed: it rises as you learn (drag the skill slider to see the climb), and the model itself is a strong teacher.
loop shape · advanced (endoreversible parameters)
How good the output gets (γ₂/η, closeness to goal) rises with checking effort, peaks, then folds back — that fold is brain-fry. Your skill — your current Dreyfus level (novice → expert) — is the ceiling it peaks at today; the amber line is the bar the stakes set. The ceiling is not fixed: it rises as you learn, and an LLM is a strong teacher — so a gap below the bar is a place to climb, not only to defer. How slowly checking pays off holds output quality low through the workslop region before the steep climb; slide how hard you check to see where a piece of work sits.
The two failures are one curve. Check too little and polished-looking output sails through — that’s the workslop region. Check everything by hand and you pass your own peak and start making more mistakes, not fewer — that’s brain-fry. In between is a zone where the checking is enough for what’s at stake and sustainable for the person doing it. Where that zone sits depends on two things you can actually measure: how good your people are at judging this kind of output today, and what the work demands.
That is what a Diagnose engagement measures — one workflow, one team, a week or two: where your people sit on this loop, what the stakes require, and where the gap is. A Build engagement then ships the evaluation suite that keeps the work in the structured zone when the model changes, the prompt changes, or the data drifts — so you know before your clients do. And because the ceiling rises as people learn, part of the answer is usually training, not tooling.