# Biostatistics Basics

> Biostatistics for FMGE Community Medicine: data types, mean median mode, SD, p-value, type I and II error, power and choosing chi-square, t-test or ANOVA.

- Canonical URL: https://prepelephant.com/topics/fmge/community-medicine/biostatistics-basics-fmge
- Exam / course: FMGE · Subject: Community Medicine
- Publisher: PrepElephant (https://prepelephant.com) — Prepared and reviewed by the PrepElephant Academic Review Team
- First published: 2026-10-02
- Last updated: 2026-10-02
- How to cite: "Biostatistics Basics", PrepElephant, https://prepelephant.com/topics/fmge/community-medicine/biostatistics-basics-fmge

## Direct answer

Data type decides the statistical test — that single sentence organises the whole chapter. Qualitative data (nominal and ordinal) is summarised by proportions and compared with the chi-square test; quantitative data is summarised by mean and standard deviation and compared with t-tests or ANOVA, switching to non-parametric tests when the distribution is skewed. Around this sit the normal distribution's 68-95-99.7 rule, standard error and confidence intervals, the null hypothesis with type I and type II errors, and power. FMGE asks tests, definitions and one-step calculations rather than derivations.

## What you must remember

- Mean is distorted by outliers; median suits skewed data (income, hospital stay); mode is the most frequent value — all three coincide in a symmetric distribution.
- The normal curve fixes 68.3% of values within ±1 SD, 95% within ±2 SD (precisely 1.96) and 99.7% within ±3 SD.
- Standard error of the mean = SD/√n; the 95% confidence interval is mean ± 1.96 SE, meaning 95 of 100 such intervals would contain the true population mean.
- The null hypothesis states no difference; p < 0.05 means a difference this large would arise by chance less than 5% of the time if the null were true.
- Type I (alpha) error rejects a true null — a false positive, conventionally 5%; type II (beta) error accepts a false null — a false negative, conventionally 20%.
- Power = 1 − beta, set at 80% or more; it rises with sample size, effect size and measurement precision — the reason sample size calculation precedes every trial.
- Parametric tests: chi-square for proportions, unpaired t-test for two independent means, paired t-test for before-after values, ANOVA for three or more means.
- Non-parametric equivalents: Mann-Whitney U, Wilcoxon signed-rank and Kruskal-Wallis; Spearman's rank correlation replaces Pearson's for skewed or ordinal data.

## Picking the right test: four worked scenarios

First: does a new antihypertensive lower systolic pressure more than the old one? Two independent groups, continuous outcome — unpaired t-test on the mean fall. Second: the same 30 patients measured before and after yoga — paired t-test, because values are paired within persons. Third: comparing HbA1c across three diets — ANOVA, because multiple pairwise t-tests inflate the type I error. Fourth: is smoking status associated with a positive treadmill test? Two categorical variables, so a chi-square on the counts. Now the twist examiners add: C-reactive protein is heavily right-skewed, so comparing CRP between groups calls for Mann-Whitney U, or log transformation before a t-test. Every choice is two questions in sequence — what type of data, and how many groups or pairings — and answering them in order selects the test.

## The SD versus SE confusion

The most persistent viva error is swapping standard deviation and standard error. Standard deviation describes variability of individual observations around the mean — patients differ from each other. Standard error describes variability of the sample mean itself across repeated samples — it shrinks as √n grows, which is why big studies have tight confidence intervals. A related trap is reading the p-value as the probability that the null hypothesis is true; it is the probability of the observed data, or data more extreme, given that the null is true. Remember also the asymmetry of the two errors: a type I error wrongly declares a drug effective (licensing a useless drug), a type II error misses an effective one; regulators fear the first more, hence alpha at a stricter 5% than beta at 20%.

## Frequently asked questions

### What is the difference between standard deviation and standard error?
Standard deviation measures scatter of individual observations around the sample mean; standard error measures how much the sample mean itself would vary across samples and equals SD divided by the square root of n.

### Which test compares two proportions, and which compares two means?
Proportions (smoking rates in two cities) are compared with the chi-square test; means of a continuous variable in two independent groups are compared with the unpaired t-test.

### How do type I and type II errors differ?
Type I error (alpha) rejects a true null — concluding a drug works when it does not; type II error (beta) fails to reject a false null — declaring an effective drug useless. Alpha is conventionally 5%, beta 20%.

### What is the power of a study and how can it be increased?
Power is 1 − beta, the ability to detect a real difference as significant, conventionally 80%; it increases with larger sample size, larger true effect size and reduced measurement variability.

### What does a 95% confidence interval of 4 to 8 mmHg for a blood-pressure difference mean?
If the study were repeated many times, 95% of such intervals would contain the true mean difference; since the interval excludes zero, the difference is significant at the 5% level.

### Which non-parametric tests replace the t-test and ANOVA?
Mann-Whitney U replaces the unpaired t-test, Wilcoxon signed-rank replaces the paired t-test, and Kruskal-Wallis replaces one-way ANOVA when data is skewed or ordinal.
