The oversight ceiling for your AI fleet.

Ten inputs from your own records. The result is the largest share of output that can go out unreviewed while your escaped-defect target still holds on the bad day, and the reviewers it takes to stay under it.

Your fleet and your reviewers

100copies
11,000

Agents that share one base model, prompt and tool set. One window is this many tasks completed at once.

10,000
100100,000
2.0%of outputs
0.5%10%

The share of outputs still defective after automated checks. Estimate it from an interleaved sample of reviewed output.

0.05ρ
0 (independent)0.5
Estimate ρ from repeated trials

Run a fixed task set through k instances. Enter the share of tasks one run passes and the share all k runs pass. The calculator inverts pass^k under the beta-binomial model.

Published τ-bench figures for one widely used model, pass^1 0.612 and pass^8 between 0.10 and 0.25, imply ρ between 0.17 and 0.42.

2.0d′, hit rate 84%
0.53.5

Measured by seeding defects into the review stream. A d′ of 2.0 catches about five defects in six; 1.0 catches about two in three.

42
1300
60
10300
12per 1,000
150
80%unreviewed
0% (review all)100%

The share of output that takes effect without a person seeing it, as your register states it now.

The defaults are the worked example in chapter 12 of The Governed Enterprise: a 100-instance fleet completing 10,000 tasks a day, reviewers who clear 60 each, and a promise of no more than 12 escaped defects per 1,000 on 95 per cent of days. Every input is a quantity you can read from your operating record or measure in a month.

Your ceiling

The ceiling35%of output can go out unreviewed
If instances failed independently43%the cost of running copies of one model
Reviewers needed at the ceiling135at 80% utilisation; you have 42, a shortfall of 93
Your ratio against the ceiling+45 ptsover the ceiling

Where the fleet stands at the ratio in force

At 80% autonomy the fleet escapes about 18 defects per 1,000 on an average day and 24 on the bad day the promise covers, against a target of 12. The 42 reviewers run at 79% utilisation.

Mean escaped, per 1,000
18.4
95th percentile, per 1,000
24.0

The mark on each bar is the target. The promise is about the tail, and a dashboard that shows only the mean will call this fleet safe.

The ceiling against correlation for a fleet of 100

Correlation between instances, ρ. The dot is your fleet; the dashed line is your ratio in force.

What moves the ceiling

LeverWhat it meansCeiling
Raise reviewer sensitivity by 0.5
Halve the correlation
Halve the residual defect rate
Accept half as many again

A whole window of this fleet carries a design effect of 5.95. An audit that reviews whole windows understates the defect-rate bound by a factor of 2.44; one that draws tasks one at a time across windows, by 1.20. Interleave the sample.

Opens a four-page report with your inputs, the ceiling, the reviewer arithmetic, the levers and the five indicators to run each quarter. Print or save it as a PDF from there.

This tool computes a ceiling from the parameters you enter, using the closed forms of a working paper that has not yet been peer reviewed. It does not constitute an audit, an assurance opinion or a statement about any particular fleet. Measuring the parameters from your own reviewed output, and setting the target and the confidence, remain your responsibility. Your inputs stay in your browser and are neither sent to Terence Kok nor reviewed by anyone.

Four known results, applied together

Review is a queue with a throughput

The result

Tasks arrive at the review function at the fleet's completion rate times the review fraction. The function has a number of reviewers and a service rate. Utilisation is the ratio, and past 0.8 the queue's waiting time grows fast enough that reviewed output is reviewed after its consequences have taken effect.

In the calculator
  • Reviewers needed at the ceiling is tasks × review fraction ÷ (reviews per reviewer × 0.8)
  • Utilisation at the ratio in force is reported beside the escaped-defect figures; above 1.0 the claimed review fraction was never achieved

Review is a detection task with a sensitivity

The result

A reviewer deciding whether an output is defective has a sensitivity d′, the separation between their response to defective and sound outputs. At the neutral criterion d′ fixes the hit rate: 2.0 catches about 84 per cent of defects, 1.0 about 69 per cent. Sensitivity falls with time on task and with low prevalence, and no change to sampling raises it.

In the calculator
  • A defect escapes when it is unreviewed, or reviewed and missed: escape share = (1 − f) + f × (1 − hit rate)
  • The first lever shows what half a point of sensitivity is worth in autonomy

Correlation widens the tail and leaves the mean alone

The result

Instances of one model share its weights, prompt, tools and blind spots, so defects arrive in correlated bursts. The expected count is unchanged; its variance is multiplied by 1 + (n − 1)ρ, the design effect. The promise is a percentile, so the ceiling is the largest autonomy ratio at which the tail of the escaped-defect distribution stays under the target, computed here with the exact beta-binomial mixture.

In the calculator
  • The second card shows the same fleet at ρ = 0; the gap is what copies cost
  • At a thousand copies and ρ = 0.4 no review fraction meets the default target: the day-to-day variation from correlation alone exceeds it at full review

The closed forms and the tables the defaults reproduce are from Terence Kok's working paper Your sample is smaller than you think: human oversight limits for fleets of identical AI agents (September 2026, not yet submitted), whose simulator and results are on the ongoing research page. The argument is chapter 12 of The Governed Enterprise, and the shorter public version is The math still caps how much you can automate. The pass^k inversion uses the τ-bench statistic; the design effect is Kish (1965); the signal-detection model is the equal-variance case. The simulation values of p, ρ and d′ are illustrative, and a fleet's own position on the curve is a measurement its deployer has yet to make.

The ceiling sizes the review function for a fleet you have already placed. To decide the fleet's tier and the controls it owes first, use theAgent Risk Assessment Matrix. To measure whether your reviewers still disagree with the system, and how their judgement holds through a shift, use theHuman-AI Interaction and Decision Quality dashboard. To test whether the board could answer for the ratio in force, use theBoard AI Oversight Checklist.

Measuring the parameters on your own fleet

The calculator does the arithmetic. Seeding the defects that measure d′, designing the interleaved sample that estimates p and ρ, and setting the target the accountable executive signs are work I do directly with clients.

Book a private session
Terence Kok
Before You Go

I built the simulator behind this page because I could not find the number anywhere else, and the first thing it told me was that the reviewers have to be sized for the bad day, and the average tells you nothing about it. Slide the correlation from zero to 0.05 and watch the ceiling fall while the mean does not move. That is the whole argument. I have not yet met a review function whose capacity had been sized before I asked, and I would like this page to change that.

Terence Kok