Tests of Significance
On this page
Direct answer
A test of significance asks whether the difference observed in a sample could plausibly have arisen by sampling fluctuation alone, given that the null hypothesis of no true difference holds; the p-value quantifies that plausibility, and comparing it with a pre-set alpha (conventionally 0.05) decides rejection. The test is chosen by data type and design: Student's t-test compares two means — one-sample, unpaired (independent) or paired — ANOVA compares three or more means, the chi-square test compares proportions on categorical data, and the Z-test serves large samples. When parametric assumptions (normality, interval-scale data) fail, their non-parametric shadows take over: Mann-Whitney U for two independent groups, Wilcoxon signed-rank for paired data, Kruskal-Wallis for several groups, Spearman's rank correlation for ranked associations. Type I error (alpha) rejects a true null; type II error (beta) accepts a false one; power, the probability of detecting a real difference, is 1 − beta and should be at least 80%.
What you must remember
- Null hypothesis machinery: H0 states no difference; the p-value is the probability of data this extreme or more if H0 is true; p below alpha (0.05) rejects H0.
- Error pair: type I (false positive, alpha, fixed at 0.05 or 0.01) versus type II (false negative, beta); power = 1 − beta, conventionally targeted at 80% or more.
- t-test family: one-sample (sample mean vs known population mean), unpaired two-sample (two independent group means), paired t-test (before-after or matched pairs on the same subjects).
- ANOVA: compares means across three or more groups through between- and within-group variance; post-hoc tests identify which pairs differ.
- Chi-square: the proportions test for categorical data laid in contingency tables; degrees of freedom = (rows − 1)(columns − 1), so a 2x2 table has df = 1.
- Yates continuity correction and Fisher exact test: when any expected cell frequency falls below 5, apply Yates correction; when expected values are very small, the Fisher exact test is the answer.
- Non-parametric map: Mann-Whitney U (unpaired t-test equivalent), Wilcoxon signed-rank (paired equivalent), Kruskal-Wallis (ANOVA equivalent), Spearman (Pearson equivalent) — for ordinal or non-normal data.
- Paired versus unpaired discipline: measuring the same subjects twice demands a paired test, accounting for within-subject correlation.
Choosing the test in an exam vignette
The stems almost always announce three things: the outcome variable's type, the number of groups, and whether observations are paired. "Mean haemoglobin of anaemic pregnant women before and after iron supplementation" — continuous outcome, two conditions, same women — is a paired t-test. "Mean systolic pressure in vegetarians versus non-vegetarians" — two independent groups — is an unpaired t-test. "Reduction in pain score across four analgesic regimens" — pain is ordinal, four groups — points to Kruskal-Wallis; had the outcome been a measured quantity, ANOVA. "Proportion of fully immunised children in two districts" — categorical outcome — chi-square on a 2x2 table with df = 1; if a cell shows an expected count under 5, Yates correction, and if it drops further, Fisher exact. "Change in knowledge score of ASHAs after training, scores heavily skewed" — Wilcoxon signed-rank.
Run the same triage in reverse to defend a protocol: power before the study, parametric checks, and the matching non-parametric fallback if they fail — the logic-chain examiners increasingly reward over pure name-matching.
How the exam frames it
The highest-yield pattern gives a two-line scenario and four test names; the discriminating words are "paired", "three groups", "proportion" and "ordinal". The second pattern tests error logic: multiple comparisons inflate type I error, hence ANOVA over repeated t-tests; increasing sample size lowers beta and raises power without touching alpha. A third asks what a p-value of 0.03 means — the probability of obtaining a result at least this extreme if the null hypothesis were true — never "the probability that the null is true", the misstatement deliberately planted among the options. Expect one memory item on degrees of freedom (chi-square 2x2: df = 1) and the parametric-to-non-parametric equivalence ladder.
Frequently asked questions
Which test compares the mean blood pressure of two independent groups?
The unpaired (independent samples) Student's t-test, assuming approximately normal distribution and comparable variances of the continuous outcome.
When is the Fisher exact test used instead of chi-square?
When categorical data yield very small expected cell frequencies (below about 5), making the chi-square approximation unreliable, particularly in 2x2 tables with small samples.
What is the non-parametric equivalent of the paired t-test?
The Wilcoxon signed-rank test, used for paired observations on ordinal scales or when differences are not normally distributed.
How do type I and type II errors differ?
Type I error rejects a true null hypothesis at rate alpha (usually 0.05); type II error fails to reject a false null at rate beta, and 1 − beta defines the study's power.
Why is ANOVA preferred over multiple pairwise t-tests?
Repeated t-tests multiply the chance of a type I error, while ANOVA tests all group means simultaneously at a single alpha level, with post-hoc tests applied only if the omnibus result is significant.