A statistics sketch · 30 September 2026
How many cases before your numbers mean anything?
When a complication is rare, a clean run feels like proof. Exact binomial arithmetic says how long that run has to be, and how often a perfectly average surgeon has a run that looks bad.
Try some numbers
Set the baseline you compare against. Everything below updates with it. The 1% starting value is only an example, not a published rate for any operation.
A teaching sketch for building intuition, not an audit tool. A real audit needs a baseline chosen for the operation and the patients, and risk adjustment for case mix.
- Observed rate
- –
- 95% interval (exact)
- –
- Could be as high as
- –
Cases needed to show you’re better
Suppose your true rate is some fraction of the baseline. How many cases give you an 80% chance that the data will show it, using a one-sided exact test at the usual 5% level? Being slightly better takes enormous numbers.
Chart values
| True rate | Cases for an 80% chance |
|---|
Bad runs happen to average surgeons
Each line is a simulated surgeon whose true rate is exactly the baseline. The line is a CUSUM: it rises after each complication, drifts down after each clean case, and sounds an alarm if it crosses the threshold. The CUSUM here is set to watch for the odds of a complication doubling. Every alarm below is false.
What this does not do
- It is not a benchmark. Choose the baseline from a source you trust for your own procedure and patients.
- It assumes every case carries the same risk. Real audits need risk adjustment for case mix. A surgeon who takes the hard cases will look worse on raw counts.
- Looking often costs something. Checking after every case and stopping at the first good-looking moment inflates false claims. The CUSUM chart shows the same effect from the other side.
- Consecutive clean cases prove less than they feel. Zero events in n cases only rules out rates above about 3/n, the “rule of three”.
Checks: the exact intervals and bounds here match SciPy’s beta-distribution values to within 10−14 on ten test cases. The case counts match an independent SciPy implementation exactly, at baselines of 1% and 5%. The CUSUM uses the standard log-likelihood-ratio weights for a Bernoulli outcome and resets to zero after an alarm.