Statistical power for programme evaluators

Stop running underpowered studies by accident

Statistics
Impact Evaluation
Monitoring & Evaluation
Author

Nichodemus Amollo

Published

November 7, 2025

Power is not a ritual number you paste into a proposal. It is a design claim: given the sample you can actually field, the variation you should expect, and the effect size that would matter for a decision, how often would a real effect be detected?

Underpowered evaluations waste partners and communities twice — once in the field, and again when a null result is over-interpreted as “no impact.”

What I want on the table before sample size is fixed

  1. Primary outcome — one main estimand, not twelve co-primary wishes
  2. Minimum effect that would change a decision — programme-relevant, not only statistically convenient
  3. Clustering — schools, facilities, villages: design effects are not optional footnotes
  4. Attrition — especially in multi-wave household work
  5. Multiple comparisons — if you will look at many outcomes, say so in the analysis plan

Practical habits

  • Compute power for the design you will run (clusters, ICC assumptions), not for an i.i.d. fantasy
  • Prefer transparent assumptions over black-box software defaults you cannot defend
  • If the affordable sample cannot detect a policy-relevant effect, redesign the question or the study — do not pretend