The oversight ceiling for your AI fleet.
Ten inputs from your own records. The result is the largest share of output that can go out unreviewed while your escaped-defect target still holds on the bad day, and the reviewers it takes to stay under it.
Your fleet and your reviewers
The defaults are the worked example in chapter 12 of The Governed Enterprise: a 100-instance fleet completing 10,000 tasks a day, reviewers who clear 60 each, and a promise of no more than 12 escaped defects per 1,000 on 95 per cent of days. Every input is a quantity you can read from your operating record or measure in a month.
Your ceiling
Where the fleet stands at the ratio in force
At 80% autonomy the fleet escapes about 18 defects per 1,000 on an average day and 24 on the bad day the promise covers, against a target of 12. The 42 reviewers run at 79% utilisation.
The mark on each bar is the target. The promise is about the tail, and a dashboard that shows only the mean will call this fleet safe.
The ceiling against correlation for a fleet of 100
Correlation between instances, ρ. The dot is your fleet; the dashed line is your ratio in force.
What moves the ceiling
| Lever | What it means | Ceiling |
|---|---|---|
| Raise reviewer sensitivity by 0.5 | ||
| Halve the correlation | ||
| Halve the residual defect rate | ||
| Accept half as many again |
A whole window of this fleet carries a design effect of 5.95. An audit that reviews whole windows understates the defect-rate bound by a factor of 2.44; one that draws tasks one at a time across windows, by 1.20. Interleave the sample.
Opens a four-page report with your inputs, the ceiling, the reviewer arithmetic, the levers and the five indicators to run each quarter. Print or save it as a PDF from there.
This tool computes a ceiling from the parameters you enter, using the closed forms of a working paper that has not yet been peer reviewed. It does not constitute an audit, an assurance opinion or a statement about any particular fleet. Measuring the parameters from your own reviewed output, and setting the target and the confidence, remain your responsibility. Your inputs stay in your browser and are neither sent to Terence Kok nor reviewed by anyone.
Four known results, applied together
Review is a queue with a throughput
Tasks arrive at the review function at the fleet's completion rate times the review fraction. The function has a number of reviewers and a service rate. Utilisation is the ratio, and past 0.8 the queue's waiting time grows fast enough that reviewed output is reviewed after its consequences have taken effect.
- Reviewers needed at the ceiling is tasks × review fraction ÷ (reviews per reviewer × 0.8)
- Utilisation at the ratio in force is reported beside the escaped-defect figures; above 1.0 the claimed review fraction was never achieved
Review is a detection task with a sensitivity
A reviewer deciding whether an output is defective has a sensitivity d′, the separation between their response to defective and sound outputs. At the neutral criterion d′ fixes the hit rate: 2.0 catches about 84 per cent of defects, 1.0 about 69 per cent. Sensitivity falls with time on task and with low prevalence, and no change to sampling raises it.
- A defect escapes when it is unreviewed, or reviewed and missed: escape share = (1 − f) + f × (1 − hit rate)
- The first lever shows what half a point of sensitivity is worth in autonomy
Correlation widens the tail and leaves the mean alone
Instances of one model share its weights, prompt, tools and blind spots, so defects arrive in correlated bursts. The expected count is unchanged; its variance is multiplied by 1 + (n − 1)ρ, the design effect. The promise is a percentile, so the ceiling is the largest autonomy ratio at which the tail of the escaped-defect distribution stays under the target, computed here with the exact beta-binomial mixture.
- The second card shows the same fleet at ρ = 0; the gap is what copies cost
- At a thousand copies and ρ = 0.4 no review fraction meets the default target: the day-to-day variation from correlation alone exceeds it at full review
The closed forms and the tables the defaults reproduce are from Terence Kok's working paper Your sample is smaller than you think: human oversight limits for fleets of identical AI agents (September 2026, not yet submitted), whose simulator and results are on the ongoing research page. The argument is chapter 12 of The Governed Enterprise, and the shorter public version is The math still caps how much you can automate. The pass^k inversion uses the τ-bench statistic; the design effect is Kish (1965); the signal-detection model is the equal-variance case. The simulation values of p, ρ and d′ are illustrative, and a fleet's own position on the curve is a measurement its deployer has yet to make.
The ceiling sizes the review function for a fleet you have already placed. To decide the fleet's tier and the controls it owes first, use theAgent Risk Assessment Matrix. To measure whether your reviewers still disagree with the system, and how their judgement holds through a shift, use theHuman-AI Interaction and Decision Quality dashboard. To test whether the board could answer for the ratio in force, use theBoard AI Oversight Checklist.
Measuring the parameters on your own fleet
The calculator does the arithmetic. Seeding the defects that measure d′, designing the interleaved sample that estimates p and ρ, and setting the target the accountable executive signs are work I do directly with clients.
Book a private session
I built the simulator behind this page because I could not find the number anywhere else, and the first thing it told me was that the reviewers have to be sized for the bad day, and the average tells you nothing about it. Slide the correlation from zero to 0.05 and watch the ceiling fall while the mean does not move. That is the whole argument. I have not yet met a review function whose capacity had been sized before I asked, and I would like this page to change that.

