how research actually works
What Does "Statistically Significant" Mean?
It does not mean big. It does not mean important. It does not mean proven. It means one narrow, technical thing — and the number people actually need is usually printed right beside it.
Statistically significant is one of the most misread phrases in science writing, and the misreading is understandable, because the everyday word significant means noteworthy. In a study it means something much smaller and much more specific: the difference between the groups is bigger than the difference chance would usually throw up on its own.
That is the entire claim. Not that the difference is large. Not that it matters to anybody. Not that it will still be there next time. This article is about what the phrase does say, what it cannot say, and which number to look at instead.

The thing being ruled out is chance
Two groups in a study never come out exactly equal. Even when the compound does absolutely nothing, one group finishes slightly ahead of the other, because the people in it were not identical, the weeks were not identical, and every measurement wobbles a little. A gap always appears. The question is never whether there is a gap.
The question is whether the gap is bigger than the ordinary wobble. Statistically significant is the label for a study answering yes.
The intuition is easiest with a coin. Flip a fair coin ten times, get six heads, and nobody raises an eyebrow — that is just how coins behave. Get nine heads out of ten and you start wondering about the coin. Nothing about the coin has changed between those two cases. What changed is how comfortably chance explains what you saw. Significance testing is that same reaction, applied to a study, with a number attached to it.
Notice what this does and does not settle. It addresses one rival explanation — the luck of who ended up in which group — and only that one. It says nothing about whether the two groups were assembled fairly, whether anyone was blinded, or whether the right thing was measured. A badly built study can produce a beautifully significant result, and it will still be a badly built study.
The p-value, in one paragraph
Here is the p-value once, in plain terms, and then we can leave the machinery alone. It is a number between zero and one that answers a hypothetical question: if the compound truly did nothing at all, how often would a study like this one produce a gap at least as big as the gap we actually got, purely through the luck of the draw? A p-value of 0.03 means such a gap would show up about three times in a hundred such studies. The smaller the number, the more strained chance becomes as an explanation. That is the whole of it.
The line below which a result gets called significant is conventionally one in twenty. It is worth knowing where that line came from, because it came from nowhere in particular: it was offered early in the twentieth century as a convenient rule of thumb, and it hardened into a rule that decides which findings get published. There is no fact of nature at that point. A result just above the line and a result just below it are not different in kind, although a great deal of publishing behaves as though they were.
The confusion around all this became serious enough that the professional body of statisticians in the United States issued a formal statement about it. Among its points: a p-value does not tell you the probability that the claim being tested is true, it does not measure the size or importance of an effect, and it should not by itself decide anything 1. A widely supported comment went further, arguing that the whole category of statistical significance should be retired, because a single threshold invites people to read a smooth scale as an on-and-off switch 5.
For a reader, the short version is this. A significant result means chance has become an awkward explanation for the gap. It does not mean the finding is true, and the awkwardness does not evaporate the moment the number creeps above the threshold.
Significant does not mean large, and it does not mean important
This is the misreading that does the most damage, and it runs in both directions.
Study forty thousand people and a difference of a hundred grams in body weight can come out significant, with a very small p-value indeed, because with numbers that large the ordinary wobble becomes tiny and even a sliver of a difference stands out against it. The result is real, detectable, and of no use to any human being. Significance measures detectability, not importance, and very large studies detect things too small to care about.
The reverse happens just as often. A difference that would matter enormously can fail to reach significance in a study of twelve people, and the paper will honestly report no significant difference. That sentence gets read as the compound does nothing. What it means is this study could not tell. The two are not close to the same statement, and the gap between them is where a great deal of confusion lives.
So there are two separate questions, and one word is being asked to answer both. Is there anything there at all? Significance speaks to that, weakly. How big is whatever is there? Significance says nothing whatsoever, and for that you need the size of the difference and the range around it.
Why small studies give noisy answers, in both directions
Everything above depends on how many people were studied, and it depends on it more heavily than most readers expect. Small studies are noisy, and the noise causes two different problems.
The first is familiar. A small study can easily miss a real effect. If a compound makes a genuine but moderate difference, a trial of twenty people will often report nothing at all, because the gap it produces sits comfortably inside the range chance produces anyway. Researchers call such a study underpowered — a study's power is simply its chance of detecting an effect that really is there. Surveys of entire research fields have found that the typical study is badly underpowered, and has been for decades 3.
The second problem is stranger and much less known. When a small study does reach significance, the effect it reports is usually too big. Think about what had to happen for it to cross the line at all: with few participants, only a large observed gap clears the threshold, so the small studies that get published are precisely the ones where the luck ran generously. The number they print then overstates whatever is truly there. This is a large part of why a striking finding from a small first study so often shrinks, or disappears, when a bigger group tries to repeat it.
Put those two problems together with the number of things researchers test and the freedom they have in choosing how to analyse them, and you have the well-known argument that a substantial share of published research findings should be expected to be false 4. Not fraudulent. False, which is a much more ordinary and much more common thing.
The more useful number, and the twenty-questions problem
There is nearly always a better number printed beside the p-value, and hardly anybody outside research looks at it. It is the confidence interval.
A confidence interval is a range: roughly, the values that are reasonably compatible with what this study found. Written out, it looks like an average loss of four kilograms, with a ninety-five per cent confidence interval of one to seven kilograms. In one line you have the study's best estimate and its honest uncertainty, and the second half is the part a p-value throws away.
Read that way, intervals answer the question significance cannot. A range of one to seven kilograms is a study that has found something, though whether it is modest or substantial is still open. A range of a fifth of a kilogram to twelve kilograms is a study that barely knows anything, whatever its p-value says. And a range running from a slight loss to a slight gain has genuinely ruled out anything dramatic in either direction — informative, useful, and reported as no significant difference. Statisticians were pressing this argument in the medical journals as far back as the nineteen-eighties 2, and it is still filtering through.
One last idea, briefly, because it explains a great many surprising headlines. Test twenty different things at the usual threshold and you should expect about one of them to come out significant by chance alone, even if nothing whatever is going on. Measure weight, sleep, mood, energy, grip strength and a dozen blood markers, and something in that list will look impressive. Nobody has cheated; it is arithmetic. The defences are for researchers to state in advance which measure they are testing, and to say plainly how many others they examined.
| What the paper says | What it means | What it does not mean |
|---|---|---|
| The result was statistically significant | Chance is an uncomfortable explanation for the gap | That the effect is big, useful, or certain to hold up |
| No significant difference was found | This study could not tell the groups apart | That the compound does nothing — a small study often cannot tell |
| A four kilogram loss, interval one to seven | Anything from modest to substantial fits the data | That four, the middle of the range, is the answer |
| Significant on one measure out of fifteen | One measure crossed the line | Much at all, unless that measure was named in advance |
None of this makes statistics untrustworthy. It makes one word untrustworthy. Significant is a narrow technical claim wearing a large everyday coat, and the useful response when you meet it is simply to look past it and ask three things: how many people were studied, how big was the difference, and how wide was the range around it. Those three answers will tell you more than the word ever could.
References
- The ASA's Statement on p-Values: Context, Process, and Purpose
- Confidence intervals rather than P values: estimation rather than hypothesis testing
- Power failure: why small sample size undermines the reliability of neuroscience
- Why Most Published Research Findings Are False
- Scientists rise up against statistical significance