Skip to main content

Educational content only. Not a substitute for medical advice. Always consult a qualified clinician.

Guides

How to read a clinical trial

Every compound page on this site references the phase-3 randomised trials that established the compound's clinical effect. This guide covers what the terms in those trial names actually mean — sample size, primary endpoint, blinding, and how effect size interacts with statistical significance.

9 minute read · Last reviewed 2026-07-13

The trial reference as a sentence in your head

When a compound page references 'STEP-1 (Wilding 2021 NEJM)' as the pivotal evidence for semaglutide 2.4 mg in obesity, that reference is a compressed sentence carrying a lot of information: the trial's name, the trial's lead author, the year of publication, and the journal. Each piece is doing work. STEP-1 tells you this was a specific pre-registered trial designed with defined endpoints (not a post-hoc analysis of pooled data). Wilding as first author tells you which group of investigators led it. NEJM 2021 tells you it was published in a top-tier peer-reviewed journal in a specific year, which is a proxy for the trial passing peer review at the level required for that journal. Understanding what to look for in a clinical trial reference is one of the more useful skills for reading this kind of catalog critically.

Sample size and statistical power

Sample size is one of the most important features of a clinical trial. A trial with 30 participants can detect only very large effects; a trial with 3000 can detect small but clinically meaningful effects. Modern phase-3 trials in metabolic disease often enrol 1000–5000 participants, which gives enough statistical power to detect clinically important differences in primary endpoints and to observe rare adverse events that would be missed in smaller trials. When a compound page cites a phase-3 trial with 2000 participants, that sample size is doing work: it means the effect estimate is precise, the confidence intervals are relatively tight, and safety signals of moderate frequency would have been detected. A study with 30 participants can support a Phase 2 exploratory conclusion but not a phase-3-level definitive claim, and reading trial references with sample size in mind clarifies what the study can and cannot support.

Primary vs secondary endpoints

Every phase-3 trial has a pre-specified primary endpoint — the outcome that the trial was statistically powered to detect. This might be percent weight loss at 68 weeks (STEP-1), HbA1c reduction at 40 weeks (SURPASS trials), or hepatic fat content reduction at 26 weeks (Falutz 2007 tesamorelin trial). The primary endpoint is what the trial is designed to prove. Secondary endpoints — other outcomes measured but not the main test — are informative but should be treated as hypothesis-generating rather than confirmatory, because trials are usually not statistically powered to definitively detect effects at secondary endpoints and the risk of false-positive findings from multiple comparisons is real. When a trial's primary endpoint is met, the trial has demonstrated what it set out to demonstrate. When only secondary endpoints are met, the finding is interesting but requires confirmation in a trial designed to test that specific hypothesis.

Blinding and randomisation

The 'randomised placebo-controlled double-blind' shorthand describes three separate design features. Randomisation means participants were assigned to the treatment arm or placebo arm by a random process, which controls for baseline differences between the two groups. Placebo-controlled means one arm received a matched inactive treatment, which controls for placebo effects and natural history. Double-blind means neither the participant nor the treating clinician knew which arm the participant was in, which controls for observer bias and behavioural changes that expectation of active treatment can produce. Trials with all three features (RCT — randomised controlled trial) are the highest-quality evidence type in clinical medicine. Trials with only some of these features (open-label, single-blind, non-randomised) give useful but weaker evidence. When a phase-3 pivotal trial has all three features and is published in a top-tier journal, the evidence quality is at the highest available level.

Effect size vs statistical significance

A trial result with a p-value below 0.05 is 'statistically significant' — the observed difference between arms is unlikely to have occurred by chance alone. But statistical significance does not automatically mean clinical significance. A very large trial can detect a very small effect with high statistical confidence, and if that small effect is not clinically meaningful, the significance is a statistical curiosity rather than a therapeutic finding. Conversely, a smaller trial might miss a clinically important effect because it wasn't powered to detect it, producing a statistically non-significant result for a compound that actually works. Reading effect size — the actual magnitude of the difference between arms — alongside statistical significance is essential. Semaglutide's STEP-1 primary endpoint (14.9% weight loss vs 2.4% for placebo, difference of 12.4 percentage points) is both statistically significant and clinically large. That combination is what makes a trial persuasive.

References

Links open external, peer-reviewed sources. Healthy Mango does not host trial data.

Continue exploring

Editorial paths through the library. Pick one and follow the trail.

Question

Other questions that touch the same biology, evidence, or laboratory concepts.