Validity and Reliability of Research Instruments
On this page
Direct answer
A bathroom scale that reads two kilograms light every morning is perfectly reliable and completely invalid — reliability is consistency, validity is truth, and a research instrument needs both before it touches a participant. Validity comes in layers: face validity (it appears to measure the concept), content validity (items cover the whole domain, judged and indexed by experts), criterion validity (scores match a gold standard, concurrently or predictively) and construct validity (scores behave as theory predicts, tested by factor analysis and contrasted groups). Reliability divides by its source of error: test-retest for stability over time, internal consistency for the interrelatedness of scale items (Cronbach's alpha, conventionally 0.70 and above), inter-rater agreement (kappa) for observation, and split-half or equivalent forms as older alternatives. Validity presupposes reliability — an inconsistent instrument cannot measure anything truthfully.
What you must remember
- Face validity: the instrument looks like it measures the concept — the weakest form, but it matters for participant acceptance.
- Content validity: items represent the entire conceptual domain; experts rate item relevance, quantified as the content validity index (CVI) — indispensable for new questionnaires.
- Criterion validity: scores correlate with a gold standard — concurrent (both measured together) or predictive (the standard comes later, as a falls-risk score predicts future falls).
- Construct validity: the instrument measures the theoretical construct — shown by convergent and discriminant evidence, factor analysis, and differences between known groups.
- Test-retest reliability: the same people measured twice, correlated — suited to stable traits; memory or genuine change can distort it.
- Internal consistency: how consistently items on one scale measure the same construct — Cronbach's alpha, commonly 0.70 or above acceptable for research tools.
- Inter-rater reliability: two observers, one phenomenon — agreement beyond chance, expressed as kappa; training and operational definitions are its foundation.
- Pilot and culture: every new or adapted tool is piloted in the study population; a translated tool needs forward and back translation and fresh evidence, because validity does not cross languages by itself.
Building a tool that earns its data
A nurse develops a 20-item knowledge questionnaire on hypoglycaemia self-management for diabetic patients. Construction: blueprint the domain first — signs, causes, treatment, danger signs — so items cover every area; write them in simple language with mixed wording. Expert review: five experts, a diabetologist, two educators and two ward nurses, rate each item for relevance; items with poor item-level CVI are revised or dropped, leaving 18. Pilot on 30 similar patients: Cronbach's alpha comes out 0.62 — two loosely related items are removed and the run repeated, reaching 0.78. Test-retest on 20 patients two weeks apart correlates 0.85 — knowledge is stable enough for the method. Construct evidence: scores differ between newly diagnosed and long-standing patients as theory predicts. Only then does the tool meet the study sample, and the dissertation's instruments chapter writes itself: development, CVI, alpha, stability, contrasted groups. The lesson examiners reward: validity and reliability are earned by procedure and reported by number, never asserted by adjective.
Where students slip
The classic viva trap is "a reliable instrument is valid" — false; the two-kilograms-light scale disproves it, while the converse, that validity implies reliability, is the true statement. Students confuse content with face validity — expert domain judgement versus mere appearance — quote an alpha with no item count or pilot detail, and forget that translation is not adaptation: an English scale rendered into Tamil without forward-back translation and a fresh pilot yields Indian data no one can interpret.
Frequently asked questions
What is the difference between validity and reliability?
Validity is whether the instrument measures what it claims; reliability is whether it measures consistently — a measure can be consistently wrong, so reliability alone never guarantees truth.
What is content validity and how is it measured?
Coverage of the entire conceptual domain by the items, judged by experts and quantified as the content validity index — indispensable for newly developed questionnaires.
What does a Cronbach's alpha of 0.78 indicate?
Acceptable internal consistency among the items of a scale — research convention sets 0.70 as the usual floor, reported with the pilot sample details.
When is test-retest reliability appropriate?
For stable traits measured twice on the same people with a time gap — a high correlation indicates stability; memory effects and genuine change limit its use.
Why do translated instruments need revalidation?
Language and culture change meaning — forward-back translation, expert review and fresh pilot testing in the new population keep the tool valid and reliable where it will be used.