Sample Size Determination

On this page
  1. Direct answer
  2. What you must remember
  3. A worked calculation from protocol to field
  4. How the exam frames it
  5. Frequently asked questions
  6. Related topics

Direct answer

Sample size grows when the difference worth detecting shrinks — that single inverse relationship organises the whole topic. For estimating a proportion, the classic formula is n = Z²pq/d², where Z is set by the confidence level (1.96 for 95%), p is the anticipated prevalence and d the precision or margin of error; for estimating a mean, n = (Z × s/d)², with s the expected standard deviation and d the tolerable error. Comparative studies add power: raising power from 80% to 90% (cutting beta) enlarges the sample, as does a smaller alpha, a more stringent one-tailed boundary, higher variability, and unequal group allocation. Cluster sampling multiplies the result by the design effect, and attrition allowances pad the final number upward. The WHO 30-cluster technique — 30 clusters, 7 children each, 210 subjects — remains the operational shorthand for immunisation coverage surveys and a favourite exam number.

What you must remember

  • Proportion formula: n = Z²pq/d²; with p = 50% (q = 50%) and d = 10% at 95% confidence, n = 96 — the most quoted worked result in PSM.
  • Mean formula: n = (Z × s/d)²; doubling the precision demanded (halving d) roughly quadruples the required sample.
  • Power input: power = 1 − beta, conventionally 80% (Z_beta 0.84); smaller expected effect sizes demand larger samples to keep power intact.
  • Alpha input: tighter alpha (0.01 versus 0.05) widens Z and enlarges n; two-tailed tests need larger n than one-tailed.
  • Variability: greater SD or a prevalence nearer 50% increases required size — pq is maximal at p = 0.5.
  • Design effect: cluster sampling imposes it (commonly taken as 2 in programme surveys); n is multiplied accordingly.
  • Attrition and non-response: inflate the calculated number (dividing by the anticipated completion fraction) before recruitment begins.
  • WHO 30-cluster survey: 30 clusters chosen by probability proportional to size, 7 children per cluster (210 total) for immunisation coverage — a number to know cold.

A worked calculation from protocol to field

A district wants to estimate full immunisation coverage expected near 50% within ±10% at 95% confidence. The formula returns n = 1.96² × 50 × 50 / 10² = 96 children. Tighten precision to ±5% and the same formula returns 384 — halving the error quadrupled the work, the inverse-square law that answers most "what happens to n" questions. Now field realities arrive: households will be sampled by clusters, so multiply by a design effect of 2 (192 at ±10%), and expect 10% non-response, so recruit about 215. Had the target been comparing coverage between two districts where a 15-percentage-point difference matters, the calculation shifts to the two-sample formula with power 80% and alpha 0.05, and the answer lands well above the estimation figure. This sequence — formula, precision choice, design effect, attrition — is exactly the reasoning an ethics committee reads in the sample-size justification, and its absence, not arithmetic error, is the commonest protocol flaw.

How the exam frames it

The dominant pattern is directional: "if precision is improved from 10% to 5%, the sample size will..." — become four times larger. The same logic covers "prevalence nearer 50%", "power increased to 90%", "alpha changed to 0.01" — all enlarge n. The second pattern is formula recall with plug-ins, nearly always the proportion formula with p = 50%; knowing that 96 answer converts the question into arithmetic. The third links to sampling methods: probability proportional to size within the 30-cluster design, stratification when subgroups must be represented, and systematic sampling with a random start. Expect one conceptual item too — an unnecessarily huge sample is not merely expensive but ethically questionable when it exposes more participants than needed to answer the question, the counterpoint that separates a statistician from a formula-user.

Frequently asked questions

What is the formula for sample size when estimating a proportion?

n = Z²pq/d², where Z reflects the confidence level, p and q the anticipated prevalence and its complement, and d the desired precision (absolute error).

What happens to sample size when the margin of error is halved?

It approximately quadruples, because n varies with the inverse square of d — the single most tested relationship in this chapter.

Why does a prevalence near 50% require the largest sample?

Because the product pq reaches its maximum at p = 0.5, maximising variability and therefore the number of observations needed for a given precision.

What is the design effect in cluster sampling?

The factor by which the sample must be inflated because observations within a cluster are correlated rather than independent; a value of about 2 is commonly assumed in programme surveys.

How does the WHO 30-cluster coverage survey work?

Thirty clusters are selected with probability proportional to size and seven eligible children sampled per cluster (210 total), giving rapid, operationally feasible immunisation coverage estimates.

Practise this in the PrepElephant app

Question banks, previous-year questions, mock tests and revision tools — for Sample Size Determination and NEET-PG Community Medicine. Free to start.

Get the free app WhatsApp