Evidence Hierarchies

Evidence Hierarchies — POPLHLTH709
1

The right design for the right question

OCEBM 2011 operationalised a key insight: “best” evidence depends on the question type, not a universal pyramid. Built for clinical questions, it is less explicit about population-level policy evaluation designs. This matrix extends that logic to include questions that the clinical framework doesn’t fully address.

How to read this. Select a question type. The designs are ordered from stronger to weaker for that specific question — not in the abstract. A design that is strong for one question type may be irrelevant or misleading for another.
Question type
Treatment & Intervention
“Does this drug/programme/service improve outcomes compared with the alternative?”

The question is causal: does the intervention cause the outcome? This requires that we can attribute differences between groups to the intervention itself, not to pre-existing differences between people who received it and those who didn’t. Randomisation is powerful here because it balances measured and — in expectation — unmeasured confounders across groups.

Design strength for this question type
↑ StrongerWeaker ↓
1
Systematic review of RCTs
Pools evidence across multiple trials, reducing random error and the influence of any single study’s quirks. Only as good as the trials included.
SR / Meta-analysis
2
Randomised controlled trial (RCT)
Random allocation creates comparable groups. Well-implemented randomisation controls confounding better than any other design — but blinding, attrition, and analysis still matter enormously.
RCT
3
Non-randomised controlled trial / quasi-experiment
Allocation is not random but there is a comparison group. Residual confounding is a major concern. Propensity scoring and regression adjustment help but cannot eliminate it.
Observational
4
Prospective cohort study
Follows people over time, measures exposure before outcome. Healthy user bias and confounding by indication are common threats when comparing treated vs untreated groups.
Observational
5
Before–after study without control group
Cannot distinguish intervention effect from natural disease course, regression to the mean, or concurrent changes. Results frequently mislead.
Weak
Important qualifications
An RCT is not automatically high certainty. Small sample, short follow-up, surrogate outcomes, industry funding, and high attrition can all reduce certainty substantially — this is what GRADE addresses (see Part 2).
Large observational studies with consistent results sometimes produce more reliable evidence than small RCTs, particularly when effect sizes are large and biological mechanism is clear.
Question type
Harm & Aetiology
“Does smoking cause lung cancer?” / “Does this vaccine cause myocarditis?”

Harm questions often cannot be answered with RCTs — especially where the exposure is strongly suspected to be harmful, randomisation is unethical or impractical. The design hierarchy shifts accordingly. Long-term and rare harms may only be detectable in large observational datasets, post-marketing surveillance, or case-control studies.

Design strength for this question type
↑ StrongerWeaker ↓
1
Systematic review of cohort or case-control studies
For harms that cannot be experimentally induced, a well-conducted SR of observational studies is the strongest available design. Heterogeneity between studies is often high.
SR
2
Large prospective cohort study
Measures exposure before outcome, can follow people over decades. The British Doctors’ Study (smoking and lung cancer) is the canonical example — transformed public health policy despite being observational.
Cohort
3
Case-control study
Efficient for rare outcomes and long-latency harms. Compares people with the outcome (cases) to those without (controls), looking back at exposure. Recall bias and control selection are the main threats.
Case-control
4
Post-marketing surveillance / passive pharmacovigilance
Essential for detecting rare vaccine or drug adverse events at population scale (e.g. VAERS, Yellow Card, CARM in NZ). Hypothesis-generating rather than hypothesis-confirming — reporting is passive and incomplete.
Surveillance
5
Case reports / case series
Critical for identifying novel harms and generating hypotheses. Cannot establish causation — no comparison group. Thalidomide and phocomelia were identified through case series: the signal matters even when the design is weak.
Signal only
Important qualifications
Rare harms are extremely unlikely to be detected in RCTs powered for efficacy. With 20,000 participants, a harm occurring in 1 in 100,000 exposures would be expected to appear only 0.2 times on average — population-scale pharmacovigilance is the only viable approach.
Bradford Hill criteria (strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, analogy) provide a structured framework for evaluating causal inference from observational harm data.
Question type
Prognosis
“What is the likely course of this disease?” / “What factors predict poor outcomes?”

Prognosis questions ask about the natural history of disease or the factors that predict outcomes. RCTs are generally not the right design here — you need to follow unselected patients over time. Inception cohort studies (beginning from first diagnosis) are the gold standard.

Design strength for this question type
↑ StrongerWeaker ↓
1
Systematic review of inception cohort studies
Synthesises evidence across multiple well-defined cohorts. Most powerful for establishing prognostic factors that generalise across settings and populations.
SR
2
Inception cohort study
Begins at disease onset (or a defined early stage), follows participants prospectively. Controls referral bias and captures the full disease trajectory. Completeness of follow-up is critical.
Cohort
3
Cohort study (non-inception)
Useful but susceptible to spectrum bias if patients are enrolled at varying disease stages. Hospital-based cohorts may over-represent severe disease.
Observational
4
Case series / case reports
May describe clinical features but cannot estimate probabilities. Survivor bias, selection effects, and absence of denominators are fundamental limitations.
Weak
Important qualifications
Loss to follow-up is the dominant threat. If sicker people drop out, estimates of bad outcomes will be too optimistic. What matters most is whether loss is differential between groups — even modest overall attrition can severely bias results if it’s concentrated in one group. Around 20% is often flagged as a red-flag threshold, but this is a heuristic, not a rule.
Treatment received during follow-up must be carefully characterised — prognosis in a modern treated population may be completely different from historical natural history data.
Question type
Diagnosis & Screening
“How accurately does this test identify people with the condition?”

Diagnostic questions require comparing a test result against a reference standard (the best available method for establishing true disease status). The key measures are sensitivity, specificity, and likelihood ratios. Cross-sectional studies with consecutive patients and a well-chosen reference standard are the appropriate design.

Design strength for this question type
↑ StrongerWeaker ↓
1
SR of cross-sectional studies (consecutive patients, consistent reference standard)
The GATE diagnostic framework: all participants receive both the test under evaluation AND the reference standard. Consecutive enrolment avoids spectrum bias.
SR
2
Cross-sectional study: consecutive patients, blind comparison to reference standard
Provides sensitivity, specificity, PPV, NPV. Blinding assessors to the reference standard result when reading the index test reduces review bias. Avoiding incorporation bias — where the index test result influences the reference standard itself — requires an independent reference standard.
Cross-sectional
3
Case-control design (diseased vs. healthy controls)
Artificially inflates sensitivity and specificity because “healthy controls” are not the same as people presenting with relevant symptoms. Often produces unrealistically optimistic estimates.
Biased
4
Expert opinion / clinical impression without reference standard
Before formal test evaluation: subjective, systematically optimistic, not reproducible.
Weak
Important qualifications
Sensitivity and specificity are not fixed properties of a test. They vary with the spectrum of disease in the population tested (spectrum effect). A test that performs well in tertiary care may perform poorly in a community screening context.
Predictive values depend on prevalence. A test with 95% sensitivity and 95% specificity applied to a population where 1% have the disease will generate more false positives than true positives. Always contextualise with local prevalence.
Question type
Prevalence & Burden
“How common is this condition in this population?” / “What proportion of people have unmet need?”

Prevalence questions are descriptive, not causal. The goal is an accurate snapshot of frequency in a defined population at a defined time. Representativeness and measurement quality matter most. An RCT usually cannot estimate population prevalence — participants are selected through eligibility criteria that distort representativeness.

Design strength for this question type
↑ StrongerWeaker ↓
1
SR of population surveys with standardised case definitions
Pools estimates across populations, enables comparison across settings. Heterogeneity in case definitions and sampling methods is the main obstacle to valid pooling.
SR
2
Population census or random-sample survey
Representative sampling, validated measures, adequate response rate. National health surveys (NZ Health Survey, BRFSS) are the standard. Census data provides denominators for rate calculations.
Survey
3
Administrative data / health records linkage
Large, complete, low cost. Captures diagnosed conditions but misses people who don’t access care. Coding errors and definitional inconsistency are common concerns. Very useful in NZ given IDI infrastructure.
Routinely collected
4
Convenience or clinic-based sample
Prevalence from hospital or clinic attenders is almost always higher than population prevalence — people seek care because they are unwell. Cannot be generalised to the community.
Biased
Important qualifications
The denominator defines the estimate. Māori and Pacific populations are frequently under-represented in health surveys due to sampling design, language barriers, or historical mistrust of government data collection. Prevalence estimates for these groups carry additional uncertainty.
Point prevalence vs period prevalence vs lifetime prevalence are not interchangeable — clarify which is reported before comparing across studies.
Question type
Policy & Programme Evaluation
“Did the sugar tax reduce obesity rates?” / “Did the smokefree legislation reduce hospitalisations?”

Policy questions are causal but randomisation is usually impossible — you cannot randomise countries or cities to policies. Quasi-experimental designs exploit natural variation in policy implementation to construct credible comparisons. These designs are largely absent from the OCEBM table.

Design strength for this question type
↑ StrongerWeaker ↓
1
SR of natural experiments and quasi-experimental studies
Synthesises evidence from multiple policy evaluations using consistent methods. Campbell Collaboration is a major source for social and policy systematic reviews; Cochrane and others also cover public health and policy topics.
SR
2
Interrupted time series (ITS)
Compares the trend before a policy change to the trend after, using the pre-period to project what would have happened without the policy. Requires sufficient data points and a clear implementation date.
Quasi-exp
3
Difference-in-differences (DiD)
Compares changes over time in an area that received a policy vs. a comparable area that did not. The “parallel trends” assumption — that the areas would have changed similarly in the absence of the policy — is the key methodological vulnerability.
Quasi-exp
4
Stepped wedge cluster RCT (randomised rollout)
All sites eventually receive the intervention, rolled out in a randomised sequence — making this a cluster RCT, not a quasi-experiment. Useful when withholding an intervention is unethical. Requires careful analysis to account for time trends; results can be sensitive to model assumptions.
Quasi-exp
5
Pre–post ecological study
Before-and-after comparison at population level without a comparison group. Cannot distinguish policy effect from concurrent trends, seasonal variation, or regression to mean. Frequently overestimates effects.
Weak
Important qualifications
RCTs of policies do exist (cluster-RCTs of workplace programmes, school-based interventions, community randomisation in LMICs) but are rare and often limited by contamination between arms and political constraints on implementation.
Implementation matters as much as effectiveness. A policy may work in one context and fail in another due to differences in enforcement, population compliance, concurrent changes, and social norms. External validity is a major challenge in policy evidence.
Population health note

The OCEBM table has no row for policy evaluation. This is not a minor omission — it is a fundamental limitation of a framework designed around individual clinical decisions. Health services planners, policy advisers, and public health practitioners routinely work with quasi-experimental evidence that the traditional hierarchy struggles to accommodate.

Question type
Equity & Distribution
“Does this intervention reduce or widen health inequities?” / “Who bears the burden of this exposure?”

Equity questions are often not asked at all in trial designs — many studies are not powered to detect differential effects by ethnicity, deprivation, or gender. Equity analysis requires disaggregated data, and finding “no significant difference” in a subgroup analysis does not mean the intervention works equally.

Design strength for this question type
↑ StrongerWeaker ↓
1
SR with explicit equity analysis using PROGRESS framework
PROGRESS: Place, Race/ethnicity/culture/language, Occupation, Gender, Religion, Education, Socioeconomic status, Social capital. Systematic disaggregation across these dimensions. Still rare in published literature.
SR
2
Kaupapa Māori / community-controlled research designs
Research designed, conducted, and governed by and for the community under study. In Aotearoa NZ, this includes data sovereignty obligations under Te Tiriti o Waitangi. Produces evidence that is legitimate and actionable for Māori communities where externally-designed studies may not.
Indigenous
3
Large cohort or administrative data with ethnicity linkage
NZ’s Integrated Data Infrastructure (IDI) enables linkage of health, social, and demographic data at individual level. Powerful for equity analysis — but requires careful attention to ethnicity data quality and prioritised ethnicity coding.
Linked data
4
Subgroup analysis from existing RCTs
Most trials are grossly underpowered for equity subgroups. Even when reported, Māori, Pacific, or low-income participants are typically too few for reliable estimates. Absence of statistical significance ≠ absence of inequity.
Often underpowered
Important qualifications
The absence of equity data is itself evidence of a system problem. When Māori and Pacific communities are excluded from studies, or when ethnicity is not collected, this reflects research priorities — not neutral gaps in knowledge.
GRADE for equity. The GRADE-Equity guidance (Welch et al. 2017) provides explicit criteria for assessing certainty of evidence about equity effects. This is not yet standard practice but is increasingly expected in WHO and NICE guidelines.
Te Tiriti o Waitangi obligations

Article 2 of Te Tiriti is widely interpreted as the basis for tino rangatiratanga — self-determination, including over data about Māori communities. Indigenous data sovereignty means Māori communities have the right to govern how data about them is collected, used, and interpreted. CARE principles (Collective benefit, Authority to control, Responsibility, Ethics) sit alongside FAIR principles in NZ public health research.

2

GRADE: from study design to certainty of evidence

Study design tells you where a piece of evidence starts in the hierarchy. GRADE tells you where it ends up — after accounting for quality, precision, consistency, directness, and publication bias across a body of studies. The two are related but not the same thing.

The traditional evidence pyramid implied a simple rule: RCT = good, observational = less good. GRADE replaced this with a more honest framework. Observational studies start at low certainty — but can be graded up. RCTs start at high certainty — but are frequently graded down. What matters is not only the design but how well it was executed, and how well its results translate to your specific question.

Certainty level
High
We are very confident that the true effect lies close to the estimate of effect. Further research is very unlikely to change our confidence in the estimate.
Certainty level
Moderate
We are moderately confident in the effect estimate. The true effect is likely to be close to the estimate, but there is a possibility that it is substantially different.
Certainty level
Low
Our confidence in the effect estimate is limited. The true effect may be substantially different from the estimate. Further research is likely to change our confidence.
Certainty level
Very Low
We have very little confidence in the effect estimate. The true effect is likely to be substantially different. Any estimate is uncertain. Action may still be warranted — but acknowledge the uncertainty.
↓ Factors that decrease certainty
R
Risk of bias Allocation concealment failed, blinding inadequate, high attrition, outcome reporting selective. Each is a route through which the estimate can be distorted.
Inconsistency Results vary widely across studies without a clear explanation. High I² in a meta-analysis is a warning sign, but unexplained clinical heterogeneity matters more than the statistic.
Indirectness The study population, intervention, comparator, or outcome differs meaningfully from your PECOT question. Surrogate outcomes (e.g. HbA1c instead of cardiovascular events) are a common form of indirectness.
~
Imprecision Wide confidence intervals that span both clinical benefit and harm. Small studies produce imprecise estimates. The 95% CI tells you the range of plausible values — if it’s wide, you’re uncertain.
P
Publication bias Positive studies are more likely to be published. Funnel plot asymmetry and systematic grey literature searching are two ways to probe for this. Industry sponsorship is associated with more favourable results across many literatures; publication and outcome reporting biases are the plausible mechanisms.
↑ Factors that increase certainty
Large effect size When relative risk is 2–5 or greater and consistent across studies, the probability that confounding alone explains the result diminishes substantially. The smoking-lung cancer association (RR ~15–20) is the canonical example.
Dose–response gradient If greater exposure produces greater effect in a consistent, monotonic pattern, this strengthens a causal interpretation even in observational data. One of Bradford Hill’s original criteria.
All plausible confounders would reduce the effect If we can argue that unmeasured confounders would attenuate rather than inflate the effect, this increases confidence that the observed effect is real. E-value calculations make this explicit.
Worked examples
How certainty can move — in both directions
Starting design
What happens to certainty
Final rating
SR of RCTs
Starts: High
−1Risk of bias: most trials industry-funded, selective outcome reporting
−1Imprecision: CI crosses line of no effect
−1Indirectness: surrogate outcome (cholesterol, not CVD events)
Low
Large cohort study
Starts: Low
+1Very large effect (RR = 15, consistent across multiple cohorts)
+1Dose–response gradient clear and monotonic
No major concerns with bias or indirectness
Moderate
Single RCT
Starts: High
−1Indirectness: trial in US tertiary care, question is NZ community setting
−1Imprecision: n=120, CI extremely wide
No inconsistency (only one study) — but also no replication
Low
ITS + natural experiment
Starts: Low
Consistent findings across multiple countries implementing same policy
+1Plausible confounders all attenuate the estimate
Direct, relevant population and outcome
Moderate

“An RCT gives you a starting point, not a conclusion. A cohort study is not automatically inferior — it depends on the question, the quality of execution, and what GRADE analysis reveals about the whole body of evidence. Certainty is earned, not assumed from design alone.

A note on guidelines and policy decisions. GRADE certainty ratings inform guideline recommendations but do not determine them. A “Strong recommendation for” can be made even on low-certainty evidence if the potential benefits are large and harms minimal (e.g. micronutrient supplementation in famines). A “Conditional recommendation against” can be appropriate even with high-certainty evidence if values and preferences vary widely. Understanding GRADE means understanding that the evidence rating and the recommendation are separate — and both require transparent reasoning.
Scroll to Top