Power is not a ritual number you paste into a proposal. It is a design claim: given the sample you can actually field, the variation you should expect, and the effect size that would matter for a decision, how often would a real effect be detected?
Underpowered evaluations waste partners and communities twice — once in the field, and again when a null result is over-interpreted as “no impact.”
What I want on the table before sample size is fixed
- Primary outcome — one main estimand, not twelve co-primary wishes
- Minimum effect that would change a decision — programme-relevant, not only statistically convenient
- Clustering — schools, facilities, villages: design effects are not optional footnotes
- Attrition — especially in multi-wave household work
- Multiple comparisons — if you will look at many outcomes, say so in the analysis plan
Practical habits
- Compute power for the design you will run (clusters, ICC assumptions), not for an i.i.d. fantasy
- Prefer transparent assumptions over black-box software defaults you cannot defend
- If the affordable sample cannot detect a policy-relevant effect, redesign the question or the study — do not pretend
Link to field systems
Power calculations assume you will obtain the sample you planned with usable data quality. That is why HFC, instrument design, and ops apps are not separate from evaluation methods: dirty or incomplete data shrink effective N.