Statistics for Nursing Research
On this page
Direct answer
A p-value of 0.03 does not mean a three per cent probability that the treatment works — it means that if the treatment truly had no effect, results this extreme would occur three times in a hundred. The working knowledge: levels of measurement (nominal, ordinal, interval, ratio) dictate which test is legitimate; descriptive statistics summarise while inferential statistics generalise; and the inferential choice follows a short decision tree — chi-square for categorical comparisons, t-tests for two group means, ANOVA for three or more, Mann-Whitney or Wilcoxon when ordinal or non-normal data break parametric rules. Type I error (rejecting a true null), type II error (missing a real effect), power at eighty per cent, and confidence intervals complete the examined toolkit.
What you must remember
- Levels of measurement: nominal (blood group), ordinal (pain score, cancer stage), interval (Celsius temperature) and ratio (weight, urine output) — the level decides the permissible statistic, never the reverse.
- Central tendency pairing: mean with standard deviation for symmetric interval or ratio data; median with interquartile range for skewed or ordinal data — a mean for a skewed variable hides the typical patient.
- The test tree: chi-square for proportions; independent t-test for two group means; paired t-test for before-and-after in the same subjects; ANOVA for three or more means; Mann-Whitney U and Wilcoxon signed-rank as their non-parametric partners; Pearson r for linear correlation, Spearman rho for ordinal.
- Type I error (alpha): rejecting a true null — a false positive, conventionally capped at 0.05; type II error (beta): missing a real effect, with power equal to 1 minus beta, conventionally 0.80.
- p-value discipline: the probability of data this extreme if the null hypothesis were true — not the probability that the null is true, and silent on effect size.
- Confidence interval: a 95 per cent interval is a range of values compatible with the data; an interval for a difference that crosses zero is not significant, however close it comes.
- Incidence versus prevalence: incidence counts new cases in a period, prevalence all existing cases at a point.
- Sampling: random, stratified, systematic, cluster or convenience — convenience limits generalisability and must be admitted.
Reading a results section without flinching
Take a typical nursing trial: two-hourly versus individualised repositioning for pressure injuries, sixty patients per arm. The results read: mean time to healing 18.4 versus 15.2 days, independent t-test p equals 0.04, 95 per cent confidence interval for the difference 0.1 to 6.3 days. The t-test fits — two independent groups, one comparison of means. The p-value of 0.04 says such a difference would be unusual if schedules truly made no difference; it does not mean the schedule works 96 per cent of the time.
The confidence interval is where clinical judgement lives: the true benefit could be as small as 0.1 days or as large as 6.3 — significant, but the lower limit asks whether the extra nursing hours are worth a third of a day. Had the study instead scored pain on a 0-10 ordinal scale with a t-test, the examiner's question is ready — ordinal data call for Mann-Whitney U.
Where candidates slip
Three slips recur. The first is the p-value catechism — "p less than 0.05 means the hypothesis is proven" — treated as an automatic deduction; the defensible sentence is about data extremity under a true null. The second is test selection by habit: t-tests on ordinal pain scores because means were computable — computability is not permissibility, and ordinal scales choose the non-parametric branch. The third is inflating statistical significance into clinical importance: a large trial can return p below 0.001 for a difference of no practical consequence. In vivas keep the error pair crisp — type I false positive, type II false negative — since "how do you reduce type II error?" (larger sample, bigger expected effect, less noise) follows directly.
Frequently asked questions
Which test compares means across three or more groups?
Analysis of variance with its F statistic, followed by post-hoc tests for specific pairs — repeated t-tests instead inflate type I error.
What does a p-value below 0.05 actually mean?
If the null hypothesis were true, data at least this extreme would occur less than five per cent of the time — evidence against the null, not the probability of the null being true.
What are type I and type II errors?
Type I rejects a true null hypothesis (false positive), controlled by alpha at typically 0.05; type II fails to reject a false null (false negative), controlled by power of at least 80 per cent.
When is the median preferred over the mean?
For skewed distributions or ordinal data — length of stay, income, pain scores — where extreme values drag the mean away from the typical experience.
How is a 95 per cent confidence interval interpreted?
As a range of values within which the true population effect plausibly lies with 95 per cent confidence; an interval spanning zero for a difference is not statistically significant.