# Statistics for Nursing Research

> Statistics for nursing research: levels of measurement, choosing chi-square t-test ANOVA, p-value meaning, confidence intervals and type I II errors.

- Canonical URL: https://prepelephant.com/topics/allied/nursing/statistics-for-nursing-research
- Exam / course: Allied Health · Subject: Nursing
- Publisher: PrepElephant (https://prepelephant.com) — Prepared and reviewed by the PrepElephant Academic Review Team
- First published: 2026-10-02
- Last updated: 2026-10-02
- How to cite: "Statistics for Nursing Research", PrepElephant, https://prepelephant.com/topics/allied/nursing/statistics-for-nursing-research

## Direct answer

A p-value of 0.03 does not mean a three per cent probability that the treatment works — it means that if the treatment truly had no effect, results this extreme would occur three times in a hundred. The working knowledge: levels of measurement (nominal, ordinal, interval, ratio) dictate which test is legitimate; descriptive statistics summarise while inferential statistics generalise; and the inferential choice follows a short decision tree — chi-square for categorical comparisons, t-tests for two group means, ANOVA for three or more, Mann-Whitney or Wilcoxon when ordinal or non-normal data break parametric rules. Type I error (rejecting a true null), type II error (missing a real effect), power at eighty per cent, and confidence intervals complete the examined toolkit.

## What you must remember

- **Levels of measurement:** nominal (blood group), ordinal (pain score, cancer stage), interval (Celsius temperature) and ratio (weight, urine output) — the level decides the permissible statistic, never the reverse.
- **Central tendency pairing:** mean with standard deviation for symmetric interval or ratio data; median with interquartile range for skewed or ordinal data — a mean for a skewed variable hides the typical patient.
- **The test tree:** chi-square for proportions; independent t-test for two group means; paired t-test for before-and-after in the same subjects; ANOVA for three or more means; Mann-Whitney U and Wilcoxon signed-rank as their non-parametric partners; Pearson r for linear correlation, Spearman rho for ordinal.
- **Type I error (alpha):** rejecting a true null — a false positive, conventionally capped at 0.05; type II error (beta): missing a real effect, with power equal to 1 minus beta, conventionally 0.80.
- **p-value discipline:** the probability of data this extreme if the null hypothesis were true — not the probability that the null is true, and silent on effect size.
- **Confidence interval:** a 95 per cent interval is a range of values compatible with the data; an interval for a difference that crosses zero is not significant, however close it comes.
- **Incidence versus prevalence:** incidence counts new cases in a period, prevalence all existing cases at a point.
- **Sampling:** random, stratified, systematic, cluster or convenience — convenience limits generalisability and must be admitted.

## Reading a results section without flinching

Take a typical nursing trial: two-hourly versus individualised repositioning for pressure injuries, sixty patients per arm. The results read: mean time to healing 18.4 versus 15.2 days, independent t-test p equals 0.04, 95 per cent confidence interval for the difference 0.1 to 6.3 days. The t-test fits — two independent groups, one comparison of means. The p-value of 0.04 says such a difference would be unusual if schedules truly made no difference; it does not mean the schedule works 96 per cent of the time.

The confidence interval is where clinical judgement lives: the true benefit could be as small as 0.1 days or as large as 6.3 — significant, but the lower limit asks whether the extra nursing hours are worth a third of a day. Had the study instead scored pain on a 0-10 ordinal scale with a t-test, the examiner's question is ready — ordinal data call for Mann-Whitney U.

## Where candidates slip

Three slips recur. The first is the p-value catechism — "p less than 0.05 means the hypothesis is proven" — treated as an automatic deduction; the defensible sentence is about data extremity under a true null. The second is test selection by habit: t-tests on ordinal pain scores because means were computable — computability is not permissibility, and ordinal scales choose the non-parametric branch. The third is inflating statistical significance into clinical importance: a large trial can return p below 0.001 for a difference of no practical consequence. In vivas keep the error pair crisp — type I false positive, type II false negative — since "how do you reduce type II error?" (larger sample, bigger expected effect, less noise) follows directly.

## Frequently asked questions

### Which test compares means across three or more groups?

Analysis of variance with its F statistic, followed by post-hoc tests for specific pairs — repeated t-tests instead inflate type I error.

### What does a p-value below 0.05 actually mean?

If the null hypothesis were true, data at least this extreme would occur less than five per cent of the time — evidence against the null, not the probability of the null being true.

### What are type I and type II errors?

Type I rejects a true null hypothesis (false positive), controlled by alpha at typically 0.05; type II fails to reject a false null (false negative), controlled by power of at least 80 per cent.

### When is the median preferred over the mean?

For skewed distributions or ordinal data — length of stay, income, pain scores — where extreme values drag the mean away from the typical experience.

### How is a 95 per cent confidence interval interpreted?

As a range of values within which the true population effect plausibly lies with 95 per cent confidence; an interval spanning zero for a difference is not statistically significant.
