how research actually works
What Is a Systematic Review?
One study is a single data point. A systematic review is the attempt to find every study ever done on a question — including the ones nobody bothered to publish — and see what they say together.
A systematic review is an attempt to find every study ever done on one question, judge each of them against rules written down beforehand, and report what they say when taken together. The word systematic refers to the search. The whole discipline of it is about finding the studies you were not hoping to find.
This is the last piece in a set about the machinery of research, and it is the one that arranges the rest. Peer review filters a single paper. A control group makes a single study interpretable. A systematic review asks the question that comes after all of that: what does the whole pile say?

Why one study is never enough
Suppose a trial is properly built. There is a control group, chance decided who went where, the assessors did not know who received what, and the result comes out clearly. Is the question settled?
No, and for a reason no amount of care can remove. A single study is one draw from a world that contains chance. If a compound genuinely does nothing, a share of the studies testing it will still produce results that look convincing. If it genuinely works, some studies will miss it entirely. That is not a flaw in those studies; it is what running one experiment means.
There is a second reason, less about luck. One study is one setting: one group of people, one dose, one length of follow-up, one laboratory's methods, one country's ordinary care. A result can be perfectly correct inside those walls and fail to travel outside them. Repeat it with older participants, or a different dose, and the answer may shift.
This is why researchers care so much about replication — the same question asked again, by different people, in a different place. When several independent groups converge on an answer, the answer starts to look like a property of the world rather than a property of one study. When they refuse to converge, something interesting is going on, and the companion article on why sources disagree is about exactly that.
How a systematic review differs from an ordinary review article
Most articles calling themselves reviews are what researchers call narrative reviews. An expert reads widely on a subject and writes an overview of it. Many are superb, and they are often the best introduction to a field that exists. The weakness is invisible from the page: nothing constrains which studies the author chose to mention. Selection rests on memory, on taste, and occasionally on interest.
A systematic review replaces that judgement with a procedure, and the procedure is written down before any searching begins.
- A protocol published in advance: the question, the planned search, the rules for including a study, the planned analysis.
- A defined search across several databases, with the exact terms recorded so that anyone else could repeat it.
- Two people screening the results independently, so one person's slip does not quietly remove a study.
- Explicit inclusion rules — which participants, which comparison, which outcome, which study designs count.
- A formal appraisal of how well each included study was run, so that weak and strong evidence are not simply added together.
- A full account of what was found, screened and rejected, with the reasons for each rejection.
That last item is why a systematic review carries a diagram of numbers near the front: so many records found, so many duplicates removed, so many screened, so many read in full, so many excluded and why, so many finally included. Those counts are the evidence that a search took place rather than a shopping trip.
The reporting standard for all of this is a checklist called PRISMA, which journals ask review authors to follow 1, and the methods themselves are set out at length in the handbook maintained by the Cochrane collaboration, which has spent decades producing reviews of this kind 5. For a reader, the practical test is blunt. If a review does not tell you how it searched and what it left out, it is a narrative review, whatever the title on it says.
Meta-analysis: putting the numbers together
A meta-analysis is the arithmetic step that sometimes follows the review: combining the results of the included studies into one overall estimate. Not every systematic review contains one, and a good review will decline to do it when the studies are too different to be added together sensibly.
The combining is not a plain average. Larger and more precise studies are given more weight than small uncertain ones, on the reasonable ground that they carry more information. Done well, the result can detect something no individual study was big enough to see on its own.
The output usually appears as a picture called a forest plot: one row per study, each with a dot for its result and a horizontal line showing its range of uncertainty, and a diamond at the bottom for the combined answer. You can read one without any statistics. A column of lines that mostly overlap is a set of studies telling a consistent story. Lines scattered on both sides of the centre means the studies disagree, and the diamond underneath is smoothing over a disagreement you can see with your own eyes.
Two cautions belong here. Pooling does not improve the ingredients: ten small, poorly controlled studies combine into one large, poorly controlled answer with deceptively tight numbers around it, which is worse than useless because it now looks authoritative. And combining studies that asked meaningfully different questions produces an average of nothing in particular — researchers call that heterogeneity, and reviewers are expected to measure it and say so.
Nor is the result final. When researchers compared meta-analyses against the large trials that were run afterwards on the same questions, the two disagreed a fair share of the time 4. A review is the best available summary of what is known. It is not a verdict.
Publication bias: the studies that never appear
Now the deepest problem in the subject, and the reason a review can be conducted impeccably and still mislead everybody.
Studies that find something are much more likely to be published than studies that find nothing. Journals prefer a discovery to a shrug. Researchers, knowing this, put the null result at the bottom of the pile and get on with the next project. Sponsors have their own reasons. Nobody has to behave badly for the outcome to follow: the published record ends up systematically more encouraging than the research that was actually carried out.
The clearest demonstration used a loophole. Trials of new medicines must be registered with the American regulator before they begin, so researchers compared the trials submitted to the regulator for a group of antidepressants against the trials that made it into journals. Of those the regulator considered positive, nearly all were published. Of those it considered negative or doubtful, most were never published at all, or appeared written up in a way that read as positive. Judging by the journals, almost every trial had succeeded. Judging by the complete record, around half had 2.
Reviewers do try to detect the hole. One method plots each study's result against its size: small studies scatter widely and large ones cluster tightly, so a complete collection makes a roughly symmetrical funnel shape. When the corner where small disappointing studies ought to sit is conspicuously empty, something is probably missing 3. It is a hint rather than proof, and it cannot recover the absent studies — only suggest that they existed.
Why reviews sit near the top, and when there is nothing to review
You will sometimes see evidence drawn as a pyramid, with opinion and individual case reports at the wide base and systematic reviews of controlled trials at the point. It is a rough ordering by how much a kind of evidence can support, and it is worth knowing even though it should not be obeyed blindly.
| Kind of evidence | What it can show | What it cannot show |
|---|---|---|
| Opinion, testimonial, single case report | That something may be worth looking into | Whether it happens more often than chance |
| Cell or animal study | What a compound does in that system | What it does in a person |
| Observational study in people | That two things travel together | That one of them causes the other |
| Randomised controlled trial | That the compound caused the difference, in that setting | Whether it holds elsewhere, or repeats |
| Systematic review of good trials | What the whole body of evidence says together | More than the studies inside it actually contain |
That final line is the important one. A systematic review of five poor trials sits at the top of the pyramid and is worth less than one large careful trial sitting below it. The level describes the shape of the evidence, not its quality, and good reviews say so plainly — often concluding that the evidence is insufficient to answer the question. That is a genuine finding, and an honest one.
Which brings us to where most research peptides actually stand. For the great majority of compounds sold for laboratory use, no systematic review exists at all. Not because reviewers have overlooked them, but because there is nothing to review: a handful of animal studies, perhaps one small early human trial, frequently nothing in people whatsoever. Reviewers occasionally publish what is called an empty review — a review that searched properly and found no eligible studies. It is a strange document, and an unusually useful one.
Absence of evidence is not evidence of absence: no review does not mean a compound does nothing. But it is also not a space that personal accounts can fill in. When there is no review behind a compound, the accurate description is that nobody yet knows, and any confident claim about it has run ahead of everything this set of articles describes.
That is the last idea in the sequence, and it is a quieter one than it looks. The useful habit is not suspicion. It is a short list of questions: who was compared with whom, how many of them, checked by whom, and repeated by whom. Most of the time you can answer all four in a couple of minutes, and most of the time those four answers are the whole story.
References
- The PRISMA 2020 statement: an updated guideline for reporting systematic reviews
- Selective publication of antidepressant trials and its influence on apparent efficacy
- Bias in meta-analysis detected by a simple, graphical test
- Discrepancies between meta-analyses and subsequent large randomized, controlled trials
- Cochrane Handbook for Systematic Reviews of Interventions