Notes & Tones · An interactive essay

Ninety-Five
per cent
confident?

A poll says the answer is 52%, give or take three, nineteen times out of twenty. You think you know what that promises. You almost certainly don't.

Scroll to begin ↓

A poll reports that 52% back the policy, give or take three points, nineteen times out of twenty. A trial finds a drug lowers blood pressure by 8 units, and is "95% confident" the true effect lies between 5 and 11. We nod along. Then, quietly, nearly everyone draws the same conclusion: that there is a 95% chance the real answer sits inside that range.

That is not what the number says. The gap between what we hear and what was actually promised is the whole subject of this essay. To close it, we have to start one step back, with an awkward fact: we almost never get to measure the thing we actually care about.

01 One sample, many possible samples

Say you want the average height of everyone in a country, or the true support for a policy across millions of voters. That true figure is the population mean, and you will essentially never see it. Measuring everyone is too slow, too dear, or simply impossible. So you take a sample, a manageable handful, and use its mean as a stand-in.

The catch is that your sample is one of countless samples you might have drawn, and each one gives a slightly different answer. That wobble is called sampling error, and at first it sounds like it should sink the whole enterprise. It doesn't, because the wobble is not chaos. It has a shape.

Below, draw samples of size n from a population and watch where their means land. Each draw drops one dot onto the pile. Change the population's shape, change the sample size, and keep drawing.

Population:
5 per sample standard error = σ/√n = 4.47 0 samples drawn
The pale curve is the population you are sampling from. The bars are where your sample means land.

Two things happen every time, whatever population you pick. The sample means pile up into a bell, centred on the true population mean, even when the population itself is lopsided or has two humps. And the bell is narrower for larger samples. This is the central limit theorem, and the width of that bell has a name and a formula: the standard error, equal to the population's standard deviation divided by the square root of the sample size. Bigger samples, tighter estimates, but only as the square root, so to halve your error you must quadruple your sample.

02 What "95% confident" really means

That predictable bell is what lets us turn a single sample into an interval. The recipe is simple: take your sample mean and reach out a fixed number of standard errors on each side, about 1.96 of them for 95%. The interval you get is built so that, across all the samples you might have drawn, 95% of the intervals it produces will contain the true mean.

Read that carefully, because it is the crux. The 95% is a property of the procedure, of the long run of intervals, not of any single interval you happen to hold. Watch the difference below. The gold line is the true mean, fixed and unmoving. Each vertical bar is one experiment's 95% interval. The blue ones catch the line; the red ones miss.

Confidence:
30 per sample
Intervals covering the true mean
0 / 0
What the recipe promises
95%
over the long run
Keep running. The share of blue intervals should hover around the confidence level you chose.

So for the one interval you actually calculated from your one real sample, the true mean is either inside it or it is not. There is no "95% chance" about that single interval; the coin has already been flipped, you just can't see the result. What you can say is that it came from a method that lands on the truth 95% of the time. Push the confidence up to 99% and the intervals grow wider to keep their promise; drop it to 90% and they tighten, but let the truth slip more often.

03 The catch: you never know σ

There is a quiet cheat in all of this. The standard error is the population's standard deviation over the square root of n, but the population's standard deviation is one more thing you almost never know. So you estimate it too, from the very same sample. Estimating your yardstick with the same data you are measuring adds a second layer of uncertainty, and for small samples that layer is not negligible.

The fix, worked out by William Gosset writing under the pen name "Student", is to stop using the normal curve and use a slightly heavier-tailed one: the t-distribution, with n minus one degrees of freedom. Shrink the sample below and watch the reference curve fatten at the tails, dragging the 95% cut-off above the familiar 1.96 and widening your interval to pay for what you don't know.

4 per sample degrees of freedom = 3 95% needs ±3.18 SE
Dashed: the normal curve you would use if you knew σ. Solid: the t-curve you must use because you don't.

For a small study the penalty is steep: with five observations the 95% interval reaches out not 1.96 but about 2.78 standard errors, roughly 40% wider. As the sample grows the extra uncertainty fades, the t-curve settles back onto the normal, and past about thirty observations the two are practically the same. This is why so much introductory statistics quietly gets away with 1.96: most textbook samples are large enough that the difference stops mattering.

04 So what?

Hold on to one sentence and you will read intervals correctly for the rest of your life: a confidence interval describes the reliability of the method, not the probability of your particular guess. The 95% lives in the long run of hypothetical repeats, not in the single range printed on the page.

From that, three habits follow. Treat "give or take three points" as the real headline, not the fine print, because it is the part that says how much to trust the rest. Don't be scandalised when a 95% interval misses; one in twenty is supposed to, and a run of studies where none ever missed would be the genuinely suspicious thing. And when an estimate looks too vague to be useful, remember the square root: the honest way to a tighter answer is almost always a bigger sample, four times the data for half the width.

A confidence interval is a promise about the net, not the fish. It tells you how often you catch the truth, never whether this cast was the one that did.