Sampling Methods
On this page
Direct answer
Surveying an entire population is rarely feasible, so epidemiology rests on sampling — selecting a fraction whose characteristics mirror the whole — and the selection method separates probability sampling (every unit has a known, non-zero chance of inclusion) from non-probability sampling (inclusion depends on convenience or judgement). The FMGE favourites are the four probability designs — simple random, systematic, stratified and cluster — plus snowball sampling for hidden populations and the WHO 30-cluster technique for immunisation coverage surveys: 30 clusters, 7 children each, 210 in total.
What you must remember
- Simple random sampling: every unit of the sampling frame (the complete list) has an equal chance, done by lottery or random number tables.
- Systematic sampling: every kth unit after a random start, where k = N/n; simple in the field but dangerous if the list has periodicity matching k.
- Stratified sampling: internally homogeneous strata (sex, district, urban or rural) are each sampled, usually proportionally; increases precision.
- Cluster sampling: natural groups (villages, schools) are sampled, then everyone or a sample within selected clusters is studied; cheap but less precise — inflating the sample by a design effect (commonly 2) compensates.
- Multistage sampling: successive steps — districts, then blocks, then villages — as in the National Family Health Survey; probability proportional to size (PPS) gives larger clusters a proportionately higher selection chance.
- Non-probability types: convenience, judgement or purposive, quota, and snowball, where existing subjects recruit peers — the method for hidden populations such as injecting drug users and female sex workers in NACO's surveillance.
- Sampling frame is the complete list from which the sample is drawn; sampling fraction is n/N.
- Sampling error (chance variation, reduced by larger samples) differs from non-sampling error (measurement and coverage mistakes, reduced by better training, not size).
Inside the 30-cluster coverage survey
The WHO expanded programme on immunisation cluster survey is the most examined design in Indian community medicine. To estimate immunisation coverage among 12-to-23-month-old children in a district, list all villages with population, then select 30 clusters with probability proportional to size: a random number chooses the first cluster and subsequent clusters follow by adding a constant sampling interval, so big villages are likelier to enter. Within each selected village, a random start household is picked and the team moves to the nearest adjacent household until 7 eligible children are found — 30 clusters × 7 = 210 children, estimating coverage within about ±10 percentage points. The design trades precision for massive cost savings: one team covers a district in days. This is also why the design effect exists — children within a village resemble each other (same anganwadi, same ANM), so each adds less new information than a randomly chosen child.
Where students slip
Stratified and cluster sampling are the classic reversal: stratification divides the population into groups that differ between themselves but are homogeneous within, sampling from every stratum; clustering samples whole natural groups and profits precisely because they are internally similar. "Stratify when strata differ from each other, cluster to save travel" is the safest hook. The second slip is systematic sampling with periodic lists — if every kth bed alternates post-operative and medical cases, the sample inherits the bias. Third, snowball sampling is often dismissed as unscientific; it is non-probability, yet the only practical route to populations without a sampling frame, which is why NACO's behavioural surveillance uses it. Finally, do not confuse sampling error with bias: a larger sample shrinks the former but can leave systematic bias untouched.
Frequently asked questions
What is the WHO 30-cluster technique and its sample size?
A cluster design estimating vaccine coverage: 30 clusters selected by probability proportional to size, 7 eligible children per cluster, 210 children in total.
What is probability proportional to size sampling?
A selection method in which larger clusters (bigger villages, for instance) have a proportionately higher probability of inclusion, keeping the overall probability of selecting any individual equal across the population.
When is snowball sampling used, with an example?
When the target population is hidden or stigmatised and no sampling frame exists — injecting drug users, female sex workers, men who have sex with men; each enrolled participant recruits peers from the network, as done in NACO's surveillance surveys.
What is the design effect in cluster sampling?
The factor by which the variance of a cluster sample exceeds that of a simple random sample of the same size, because individuals within a cluster are alike; a design effect of 2 is the customary assumption, so cluster samples are doubled.
How is systematic sampling performed and what is its main pitfall?
Choose a random start between 1 and k, then take every kth unit (k = population ÷ sample size); the pitfall is a periodic list whose rhythm coincides with k.
What is the difference between sampling error and non-sampling error?
Sampling error is the chance difference between sample and population that falls as sample size rises; non-sampling error arises from faulty measurement, non-response or wrong coverage and cannot be fixed by enlarging the sample.