how research actually works
What Is a Control Group?
Evidence is not a measurement. It is a comparison — and the comparison is the half that claims most often leave out, usually without anyone noticing that it is missing.
A control group is a second set of participants in the same study who do not receive the thing being tested. Everything else about them is kept as similar as possible: the same weeks, the same clinics, the same measurements, the same attention. They exist to answer the one question that makes any result mean something — compared with what?
The idea is plain enough that it hardly seems to need explaining. It needs explaining because the commonest way claims mislead people is not fraud, and it is not clever statistics. It is a missing comparison, and once you can see the gap you cannot stop seeing it.

Compared with what?
Start with a number. Eighty-four per cent of participants improved. It sounds substantial, and on its own it is empty. Improved compared with what? With how they felt at the start, when things were bad enough to make them volunteer? With people who took nothing? With people who took the treatment that already exists? Those are three different claims, and only the last two say anything about the compound.
So the control group is not one feature of a study among many. It is the study. The blinding, the randomising, the statistics at the end — all of it is machinery built to make a single comparison trustworthy. Take the comparison away and there is nothing left for the machinery to protect.
What the control group receives depends on the question being asked. It might be nothing. It might be a placebo, an inert imitation of the treatment, which has an article of its own. It might be the established treatment, when leaving people with nothing would be wrong. Each choice turns the study into a different question, and it is worth noticing which one you are being shown.
The comparison also has to happen at the same time as the thing it is being compared with. Same months, same recruiting, same rules about who can join, same instruments doing the measuring. A comparison assembled later, or borrowed from elsewhere, brings differences with it that have nothing to do with the treatment.
Why before-and-after in one group is not a controlled result
Here is the design you will meet most often, in and out of medicine. Take a group of people. Measure something. Give them the compound for twelve weeks. Measure the same thing again. Report the change.
It feels rigorous. There are real measurements, taken twice, often with a statistical test at the end to make it official. But there is no control group, and a group's own starting value is not a control. By the second measurement everyone is twelve weeks older, the season has changed, they have been watched and encouraged and asked how they are getting on, and they knew from the first day exactly what they were taking.
One effect deserves its own name because it manufactures impressive numbers all by itself. It is called regression to the mean, and the plain version is this: people come forward when things are bad. The flare-up week, the terrible month of sleep, the blood result that came back high — that is when anyone goes looking for help. Anything that fluctuates keeps fluctuating, and the measurement after an unusually bad one tends to be less bad. Measure people at their worst and measure them again later, and you will record an improvement in a group that received nothing at all.
There is a halfway house that looks more respectable: comparing the treated group with an earlier one — patients seen last year, or numbers lifted from a previous study. These are called historical controls, and they are known to flatter. A review that lined up trials of the same treatments found that those using historical controls concluded the treatment worked far more often than the randomised trials of the very same treatments did 4. No dishonesty is required. The earlier patients were diagnosed with older methods, cared for by older routines, and chosen by different criteria.
Confounding, with an example that has nothing to do with medicine
Confounding is the technical name for the underlying trap. A confounder is a third thing that causes both of the things you are looking at, so the two appear connected although neither one causes the other.
The cleanest example involves no medicine whatsoever. Take a large group of primary school children and measure two things: shoe size and reading ability. The two turn out to be strongly related. Children with bigger feet read better. The finding is entirely real, and you can repeat it in any school you like.
Nobody imagines that bigger shoes cause reading. The hidden third thing is age. Seven-year-olds have bigger feet than five-year-olds and they also read better, so age drives both measurements at once. Compare children who are all the same age and the relationship vanishes. Age is the confounder, and the correlation was never a clue about feet.
In real research the third thing is rarely as obvious as age. Suppose people who take a particular supplement turn out, on average, to be healthier than people who do not. Is the supplement doing it? Or do people who buy supplements also tend to sleep more, drink less, exercise more, earn more and see a doctor sooner? Any one of those produces better health on its own, and all of them travel together.
Researchers can correct for confounders they have measured, comparing like with like using statistics. The permanent difficulty is the ones nobody thought of, or could not measure. You cannot adjust for a factor you do not know exists — which is precisely the problem the next section solves, and it solves it in a way that sounds too simple to work.
Randomisation, and what it protects against
Randomisation means chance decides which group each person goes into — a coin flip, or a computer doing the equivalent. Nobody chooses. It is the least intuitive part of study design and probably the most valuable single thing in it.
Picture researchers assigning people themselves, with nothing but good intentions. This patient looks frail, so she goes in the gentler arm. That one seems unlikely to keep his appointments, so he goes in the control arm where it will matter less. Every decision is defensible on its own, and the two groups now differ before anybody has been given anything. From that point on, the compound and the sorting cannot be told apart.
Chance has no opinions. Across enough people it spreads the frail and the fit, the determined and the half-hearted, evenly between the groups — and it does the same for every factor nobody thought to write down. That is the property that makes it irreplaceable. Statistical adjustment handles the confounders you measured. Randomisation handles the ones you never imagined.
It has to be real chance, though. Assigning by alternate patients, by day of the week, or by odd and even birth dates is not random, because it can be predicted, and if the person enrolling participants can predict what comes next, they can wait for a better-suited patient before signing the next one up 2. Keeping the upcoming assignments hidden from whoever recruits is called allocation concealment, and it is the practical half of the idea.
This is not a theoretical worry. When published trials were sorted by how well the allocation had been concealed, the ones that did it inadequately reported treatment effects around a third larger than the ones that did it properly 3. Same treatments, more flattering answers, produced by nothing except how people were assigned.
The design is also newer than most people assume. The study usually described as the first properly randomised trial in medicine appeared in 1948, testing streptomycin against tuberculosis, and its authors explained plainly why chance had been used to allocate patients 1. Nearly everything now taken for granted about evidence starts there, well within living memory.
What a well-matched control looks like
Put together, a good comparison group is thoroughly unglamorous. There is nothing clever in it at all.
- Recruited at the same time, from the same places, under the same rules about who may join.
- Assigned by chance, with the sequence hidden from whoever enrols people.
- Followed on the same schedule, seen as often, and asked the same questions.
- Measured with the same instruments, ideally by assessors who do not know who is in which group.
- Given the same package of attention — a placebo where that is possible, the established treatment where withholding one would be wrong.
- Counted at the end even if they dropped out, so that a group cannot improve simply by losing its worst cases.
That final item is quietly important. If the people who felt no benefit drift away and are left out of the analysis, the group that remains looks better without anything having actually changed. The fix has a name you will see in methods sections: analysis by intention to treat, meaning everyone is counted in the group they were assigned to, whatever happened afterwards.
None of this asks you to be able to run a trial. It gives you one question to carry around, and it is the same question every time. Compared with what? If a claim cannot answer it — because there was no second group, or the second group was last year's patients, or the only comparison offered is how people felt before they started — then something was certainly measured, but the compound has not been shown to be the cause of it.
References
- Streptomycin treatment of pulmonary tuberculosis
- Generation of allocation sequences in randomised trials: chance, not choice
- Empirical evidence of bias. Dimensions of methodological quality associated with estimates of treatment effects in controlled trials
- Randomized versus historical controls for clinical trials