# Validity and Reliability of Research Instruments

> Validity and reliability of Nursing research instruments: face, content, criterion and construct validity, Cronbach's alpha, test-retest and kappa.

- Canonical URL: https://prepelephant.com/topics/allied/nursing/instruments-validity-nursing
- Exam / course: Allied Health · Subject: Nursing
- Publisher: PrepElephant (https://prepelephant.com) — Prepared and reviewed by the PrepElephant Academic Review Team
- First published: 2026-10-02
- Last updated: 2026-10-02
- How to cite: "Validity and Reliability of Research Instruments", PrepElephant, https://prepelephant.com/topics/allied/nursing/instruments-validity-nursing

## Direct answer

A bathroom scale that reads two kilograms light every morning is perfectly reliable and completely invalid — reliability is consistency, validity is truth, and a research instrument needs both before it touches a participant. Validity comes in layers: face validity (it appears to measure the concept), content validity (items cover the whole domain, judged and indexed by experts), criterion validity (scores match a gold standard, concurrently or predictively) and construct validity (scores behave as theory predicts, tested by factor analysis and contrasted groups). Reliability divides by its source of error: test-retest for stability over time, internal consistency for the interrelatedness of scale items (Cronbach's alpha, conventionally 0.70 and above), inter-rater agreement (kappa) for observation, and split-half or equivalent forms as older alternatives. Validity presupposes reliability — an inconsistent instrument cannot measure anything truthfully.

## What you must remember

- **Face validity:** the instrument looks like it measures the concept — the weakest form, but it matters for participant acceptance.
- **Content validity:** items represent the entire conceptual domain; experts rate item relevance, quantified as the content validity index (CVI) — indispensable for new questionnaires.
- **Criterion validity:** scores correlate with a gold standard — concurrent (both measured together) or predictive (the standard comes later, as a falls-risk score predicts future falls).
- **Construct validity:** the instrument measures the theoretical construct — shown by convergent and discriminant evidence, factor analysis, and differences between known groups.
- **Test-retest reliability:** the same people measured twice, correlated — suited to stable traits; memory or genuine change can distort it.
- **Internal consistency:** how consistently items on one scale measure the same construct — Cronbach's alpha, commonly 0.70 or above acceptable for research tools.
- **Inter-rater reliability:** two observers, one phenomenon — agreement beyond chance, expressed as kappa; training and operational definitions are its foundation.
- **Pilot and culture:** every new or adapted tool is piloted in the study population; a translated tool needs forward and back translation and fresh evidence, because validity does not cross languages by itself.

## Building a tool that earns its data

A nurse develops a 20-item knowledge questionnaire on hypoglycaemia self-management for diabetic patients. Construction: blueprint the domain first — signs, causes, treatment, danger signs — so items cover every area; write them in simple language with mixed wording. Expert review: five experts, a diabetologist, two educators and two ward nurses, rate each item for relevance; items with poor item-level CVI are revised or dropped, leaving 18. Pilot on 30 similar patients: Cronbach's alpha comes out 0.62 — two loosely related items are removed and the run repeated, reaching 0.78. Test-retest on 20 patients two weeks apart correlates 0.85 — knowledge is stable enough for the method. Construct evidence: scores differ between newly diagnosed and long-standing patients as theory predicts. Only then does the tool meet the study sample, and the dissertation's instruments chapter writes itself: development, CVI, alpha, stability, contrasted groups. The lesson examiners reward: validity and reliability are earned by procedure and reported by number, never asserted by adjective.

## Where students slip

The classic viva trap is "a reliable instrument is valid" — false; the two-kilograms-light scale disproves it, while the converse, that validity implies reliability, is the true statement. Students confuse content with face validity — expert domain judgement versus mere appearance — quote an alpha with no item count or pilot detail, and forget that translation is not adaptation: an English scale rendered into Tamil without forward-back translation and a fresh pilot yields Indian data no one can interpret.

## Frequently asked questions

### What is the difference between validity and reliability?

Validity is whether the instrument measures what it claims; reliability is whether it measures consistently — a measure can be consistently wrong, so reliability alone never guarantees truth.

### What is content validity and how is it measured?

Coverage of the entire conceptual domain by the items, judged by experts and quantified as the content validity index — indispensable for newly developed questionnaires.

### What does a Cronbach's alpha of 0.78 indicate?

Acceptable internal consistency among the items of a scale — research convention sets 0.70 as the usual floor, reported with the pilot sample details.

### When is test-retest reliability appropriate?

For stable traits measured twice on the same people with a time gap — a high correlation indicates stability; memory effects and genuine change limit its use.

### Why do translated instruments need revalidation?

Language and culture change meaning — forward-back translation, expert review and fresh pilot testing in the new population keep the tool valid and reliable where it will be used.
