Biostatistics for BDS
On this page
Direct answer
Statistics first summarises and then infers: descriptive statistics condense a sample into a mean, median or mode with a measure of spread such as the standard deviation, while inferential statistics test whether an observed difference is more than chance could explain. The 95 per cent confidence interval — mean plus or minus 1.96 times the standard error of the mean, where SEM = SD/square root of n — is the modern way to report precision. Choosing the right test follows two questions: what type of data (qualitative categories versus quantitative measurements) and how many groups or measurements are being compared. Chi-square handles proportions, the t-test compares two means, ANOVA compares three or more, and when data are skewed or ordinal — as dental indices often are — the non-parametric equivalents (Mann-Whitney, Wilcoxon signed-rank, Kruskal-Wallis) take over.
What you must remember
- Central tendency matches data shape: mean for symmetric quantitative data, median for skewed data and ordinals, mode for nominal data; DMFT distributions are typically right-skewed, which is why medians often serve better than means.
- Dispersion set: range, interquartile range, variance and standard deviation; 68-95-99.7 per cent of a normal distribution lies within one, two and three SDs respectively.
- SEM versus SD: SD describes the spread of individuals; SEM (SD/square root n) describes the spread of sample means and shrinks as samples grow — mixing them up is the commonest reporting error.
- Qualitative data tests: chi-square for comparing proportions across groups (with Yates' correction for 2x2 tables), Fisher's exact test when expected cell counts fall below five, McNemar's test for paired proportions.
- Quantitative ladder: unpaired t-test for two independent means, paired t-test for before-after measurements on the same subjects, one-way ANOVA (with post-hoc tests) for three or more groups.
- Non-parametric mirrors: Mann-Whitney U (two independent groups), Wilcoxon signed-rank (paired), Kruskal-Wallis (three or more groups), Spearman's rank correlation for non-normal pairs.
- Association measures: Pearson's correlation coefficient r runs from -1 to +1 and measures linear association, never agreement or causation; regression goes further and predicts.
- The p-value: the probability of results this extreme if the null hypothesis were true — not the probability that the null hypothesis is true; conventional significance is p below 0.05.
Matching tests to five dental datasets
Walk through a thesis clinic's week. Monday, compare plaque scores (ordinal scale) between a test and a control toothpaste group: Mann-Whitney U, because ordinal data violate normality. Tuesday, compare DMFT means of boys versus girls in one school: unpaired t-test, provided the distributions pass a normality check — otherwise Mann-Whitney again. Wednesday, compare gingival index before and after supervised brushing in the same children: paired t-test, since each child is their own control. Thursday, compare DMFT across three fluoride-content zones of a district: one-way ANOVA, then a post-hoc test to find which zones actually differ. Friday, compare the proportion of children with fluorosis between two villages: chi-square on the 2x2 table. Notice the discipline in each step — data type first, number of groups second, independence versus pairing third — and the test names follow automatically. One more habit from this week: report the confidence interval alongside every p-value, because p = 0.04 on a tiny sample can hide a useless estimate.
Where students slip in statistics
The recurring viva failure is SD-SEM confusion: a journal line reading "DMFT 2.1 +/- 0.3" is interpretable only if you know whether 0.3 is the spread of children or the wobble of the mean. The second is applying a t-test to DMFT data without checking skew — caries counts cluster at low values with a long tail of high-caries children, which is why many examiners accept non-parametric handling. Third, students say "p = 0.06, so there is no difference" — the correct phrasing is that the study failed to demonstrate a difference, often a power problem rather than proof of equivalence. Fourth, a significant chi-square says groups differ in proportion but says nothing about why; Finally, correlation r = 0.9 does not mean agreement — two examiners can correlate perfectly while one scores everyone half a unit higher, which is exactly why calibration studies report kappa statistics instead.
Frequently asked questions
Which test compares two independent group means?
The unpaired (two-sample) t-test, assuming approximately normal distributions and comparable variances; otherwise use the Mann-Whitney U test.
When is ANOVA used instead of repeated t-tests?
When comparing means across three or more groups, because repeated t-tests inflate the overall type I error; post-hoc tests then locate the specific differences.
What does a 95 per cent confidence interval mean?
A range constructed so that if the study were repeated many times, 95 per cent of such intervals would contain the true population value; it expresses precision, unlike the p-value alone.
Why is the paired t-test used in before-after studies?
Because the same subjects are measured twice, so each person acts as their own control, removing between-person variability from the comparison.
What is the non-parametric equivalent of ANOVA?
The Kruskal-Wallis test, used when data are ordinal or non-normally distributed across three or more independent groups.