Tests of Significance

On this page
  1. Direct answer
  2. What you must remember
  3. Choosing the test in an exam vignette
  4. How the exam frames it
  5. Frequently asked questions
  6. Related topics

Direct answer

A test of significance asks whether the difference observed in a sample could plausibly have arisen by sampling fluctuation alone, given that the null hypothesis of no true difference holds; the p-value quantifies that plausibility, and comparing it with a pre-set alpha (conventionally 0.05) decides rejection. The test is chosen by data type and design: Student's t-test compares two means — one-sample, unpaired (independent) or paired — ANOVA compares three or more means, the chi-square test compares proportions on categorical data, and the Z-test serves large samples. When parametric assumptions (normality, interval-scale data) fail, their non-parametric shadows take over: Mann-Whitney U for two independent groups, Wilcoxon signed-rank for paired data, Kruskal-Wallis for several groups, Spearman's rank correlation for ranked associations. Type I error (alpha) rejects a true null; type II error (beta) accepts a false one; power, the probability of detecting a real difference, is 1 − beta and should be at least 80%.

What you must remember

  • Null hypothesis machinery: H0 states no difference; the p-value is the probability of data this extreme or more if H0 is true; p below alpha (0.05) rejects H0.
  • Error pair: type I (false positive, alpha, fixed at 0.05 or 0.01) versus type II (false negative, beta); power = 1 − beta, conventionally targeted at 80% or more.
  • t-test family: one-sample (sample mean vs known population mean), unpaired two-sample (two independent group means), paired t-test (before-after or matched pairs on the same subjects).
  • ANOVA: compares means across three or more groups through between- and within-group variance; post-hoc tests identify which pairs differ.
  • Chi-square: the proportions test for categorical data laid in contingency tables; degrees of freedom = (rows − 1)(columns − 1), so a 2x2 table has df = 1.
  • Yates continuity correction and Fisher exact test: when any expected cell frequency falls below 5, apply Yates correction; when expected values are very small, the Fisher exact test is the answer.
  • Non-parametric map: Mann-Whitney U (unpaired t-test equivalent), Wilcoxon signed-rank (paired equivalent), Kruskal-Wallis (ANOVA equivalent), Spearman (Pearson equivalent) — for ordinal or non-normal data.
  • Paired versus unpaired discipline: measuring the same subjects twice demands a paired test, accounting for within-subject correlation.

Choosing the test in an exam vignette

The stems almost always announce three things: the outcome variable's type, the number of groups, and whether observations are paired. "Mean haemoglobin of anaemic pregnant women before and after iron supplementation" — continuous outcome, two conditions, same women — is a paired t-test. "Mean systolic pressure in vegetarians versus non-vegetarians" — two independent groups — is an unpaired t-test. "Reduction in pain score across four analgesic regimens" — pain is ordinal, four groups — points to Kruskal-Wallis; had the outcome been a measured quantity, ANOVA. "Proportion of fully immunised children in two districts" — categorical outcome — chi-square on a 2x2 table with df = 1; if a cell shows an expected count under 5, Yates correction, and if it drops further, Fisher exact. "Change in knowledge score of ASHAs after training, scores heavily skewed" — Wilcoxon signed-rank.

Run the same triage in reverse to defend a protocol: power before the study, parametric checks, and the matching non-parametric fallback if they fail — the logic-chain examiners increasingly reward over pure name-matching.

How the exam frames it

The highest-yield pattern gives a two-line scenario and four test names; the discriminating words are "paired", "three groups", "proportion" and "ordinal". The second pattern tests error logic: multiple comparisons inflate type I error, hence ANOVA over repeated t-tests; increasing sample size lowers beta and raises power without touching alpha. A third asks what a p-value of 0.03 means — the probability of obtaining a result at least this extreme if the null hypothesis were true — never "the probability that the null is true", the misstatement deliberately planted among the options. Expect one memory item on degrees of freedom (chi-square 2x2: df = 1) and the parametric-to-non-parametric equivalence ladder.

Frequently asked questions

Which test compares the mean blood pressure of two independent groups?

The unpaired (independent samples) Student's t-test, assuming approximately normal distribution and comparable variances of the continuous outcome.

When is the Fisher exact test used instead of chi-square?

When categorical data yield very small expected cell frequencies (below about 5), making the chi-square approximation unreliable, particularly in 2x2 tables with small samples.

What is the non-parametric equivalent of the paired t-test?

The Wilcoxon signed-rank test, used for paired observations on ordinal scales or when differences are not normally distributed.

How do type I and type II errors differ?

Type I error rejects a true null hypothesis at rate alpha (usually 0.05); type II error fails to reject a false null at rate beta, and 1 − beta defines the study's power.

Why is ANOVA preferred over multiple pairwise t-tests?

Repeated t-tests multiply the chance of a type I error, while ANOVA tests all group means simultaneously at a single alpha level, with post-hoc tests applied only if the omnibus result is significant.

Same topic for other exams

Practise this in the PrepElephant app

Question banks, previous-year questions, mock tests and revision tools — for Tests of Significance and NEET-PG Community Medicine. Free to start.

Get the free app WhatsApp