Not all evidence is equal.
But equal for what?
Study design hierarchies exist because some designs are better at controlling error than others — but only for particular question types. And even the best-designed study can produce uncertain evidence. This page works through both ideas.
The right design for the right question
OCEBM 2011 operationalised a key insight: “best” evidence depends on the question type, not a universal pyramid. Built for clinical questions, it is less explicit about population-level policy evaluation designs. This matrix extends that logic to include questions that the clinical framework doesn’t fully address.
The OCEBM table has no row for policy evaluation. This is not a minor omission — it is a fundamental limitation of a framework designed around individual clinical decisions. Health services planners, policy advisers, and public health practitioners routinely work with quasi-experimental evidence that the traditional hierarchy struggles to accommodate.
Article 2 of Te Tiriti is widely interpreted as the basis for tino rangatiratanga — self-determination, including over data about Māori communities. Indigenous data sovereignty means Māori communities have the right to govern how data about them is collected, used, and interpreted. CARE principles (Collective benefit, Authority to control, Responsibility, Ethics) sit alongside FAIR principles in NZ public health research.
GRADE: from study design to certainty of evidence
Study design tells you where a piece of evidence starts in the hierarchy. GRADE tells you where it ends up — after accounting for quality, precision, consistency, directness, and publication bias across a body of studies. The two are related but not the same thing.
The traditional evidence pyramid implied a simple rule: RCT = good, observational = less good. GRADE replaced this with a more honest framework. Observational studies start at low certainty — but can be graded up. RCTs start at high certainty — but are frequently graded down. What matters is not only the design but how well it was executed, and how well its results translate to your specific question.
−1Imprecision: CI crosses line of no effect
−1Indirectness: surrogate outcome (cholesterol, not CVD events)
+1Dose–response gradient clear and monotonic
No major concerns with bias or indirectness
−1Imprecision: n=120, CI extremely wide
No inconsistency (only one study) — but also no replication
+1Plausible confounders all attenuate the estimate
Direct, relevant population and outcome
“An RCT gives you a starting point, not a conclusion. A cohort study is not automatically inferior — it depends on the question, the quality of execution, and what GRADE analysis reveals about the whole body of evidence. Certainty is earned, not assumed from design alone.“