Correlation and Regression
On this page
Direct answer
Correlation measures the strength and direction of a linear relationship between two quantitative variables through a single coefficient: Pearson's r for normally distributed continuous data, Spearman's rank correlation for ordinal or non-normal data, both running from −1 (perfect negative) through 0 (no linear relation) to +1 (perfect positive). Regression goes further and draws the line: simple linear regression fits Y = a + bX by least squares, so that b estimates how much Y changes per unit of X, and Y can be predicted from X. The coefficient of determination, r-squared, is the share of Y's variance explained by X — an r of 0.8 means 64% — and logistic regression extends the idea to binary outcomes such as disease yes/no, yielding adjusted odds ratios for each predictor. Correlation, however strong, never proves causation: that inference needs design, not arithmetic.
What you must remember
- Pearson's r: parametric, for two continuous, roughly normal variables; sign gives direction, magnitude gives strength; r is unitless.
- Spearman's rho: rank-based non-parametric alternative for ordinal or skewed data and for outlier-laden series.
- Interpretation bands: 0 no linear relationship; below 0.3 weak, 0.3-0.7 moderate, above 0.7 strong — with the caveat that r detects only linear patterning.
- r-squared: variance explained; r = 0.8 → r² = 64%; the companion statistic to any correlation quoted in a paper.
- Simple linear regression: Y = a + bX; b (regression coefficient) is the change in Y per unit X; a is the intercept; fitted by least squares, minimising squared residuals.
- Multiple regression: several predictors of a continuous outcome; logistic regression for binary outcomes reports adjusted odds ratios, the workhorse of clinical research tables.
- Correlation is not causation: confounding, reverse causation and coincidence all generate strong r; only an experiment or rigorously designed cohort can argue causality.
- Scatter first: always inspect the scatter diagram — a strong curvilinear relationship (drug effect versus dose) can yield r near 0 and still be a real relationship.
Worked example: birth weight and gestational age
Plot 200 newborns with gestational age on the x-axis and birth weight on the y-axis: the cloud rises from lower left to upper right. Pearson's r computes to 0.8 — strong positive — and r-squared says gestational age explains about 64% of birth-weight variation; the remaining 36% belongs to maternal nutrition, sex, parity and measurement noise. Now fit the regression: Y = −2,500 + 165X (weight in grams, age in weeks), meaning each extra week in utero adds about 165 g, on average, within the observed range — the phrase "within the observed range" matters, because extrapolating to 60 weeks is meaningless. If instead the outcome were low birth weight (yes/no) predicted from maternal haemoglobin, weight and smoking, the tool changes to multiple logistic regression, and each variable reports an adjusted odds ratio with its confidence interval, holding the others constant — which is precisely how confounders are tamed in analysis.
The teaching sequence in these two paragraphs is the one to reproduce in short answers: plot, quantify (r, r²), model (b, or adjusted OR), interpret within range, and never let the arithmetic outrun the design.
How the exam frames it
Three recurring items. First, the value-of-r questions: r = −0.9 describes a strong inverse relationship, and candidates who miss the sign miss the mark; r = 0 with a U-shaped scatter tests whether you know r measures only linear association. Second, the correlation-versus-regression distinction: correlation is symmetric (r of X-with-Y equals Y-with-X) while regression is not — regressing Y on X yields a different line from X on Y — and prediction belongs to regression alone. Third, the causation trap: a vignette correlating television ownership with coronary deaths expects "confounding by affluence", not "televisions cause heart disease". The r-to-r-squared conversion (0.8 → 64%, 0.5 → 25%) is the cheapest mark in the chapter and is asked nearly every year in some form.
Frequently asked questions
What does a Pearson correlation coefficient of −0.85 indicate?
A strong inverse linear relationship between two continuous variables — as one rises, the other falls, with about 72% of variance shared (r² = 0.72).
When should Spearman's rank correlation be used instead of Pearson's?
For ordinal data, non-normally distributed variables, or series with influential outliers, since ranking blunts the effect of extreme values.
What does the regression coefficient b represent in Y = a + bX?
The average change in Y for a one-unit increase in X, holding the model's assumptions; in multiple regression each coefficient is adjusted for the other variables.
Why can correlation never establish causation?
Because association may reflect confounding, reverse causation or chance, and only an experimental or well-designed longitudinal design can support a causal claim.
Which regression handles a binary outcome such as disease present or absent?
Logistic regression, which models the log-odds of the outcome and expresses each predictor's effect as an adjusted odds ratio.