# Correlation and Regression

> Correlation and regression for NEET-PG Community Medicine: Pearson and Spearman coefficients, r-squared, simple linear and logistic regression explained.

- Canonical URL: https://prepelephant.com/topics/neet-pg/community-medicine/correlation-and-regression
- Exam / course: NEET-PG · Subject: Community Medicine
- Publisher: PrepElephant (https://prepelephant.com) — Prepared and reviewed by the PrepElephant Academic Review Team
- First published: 2026-10-02
- Last updated: 2026-10-02
- How to cite: "Correlation and Regression", PrepElephant, https://prepelephant.com/topics/neet-pg/community-medicine/correlation-and-regression

## Direct answer

Correlation measures the strength and direction of a linear relationship between two quantitative variables through a single coefficient: Pearson's r for normally distributed continuous data, Spearman's rank correlation for ordinal or non-normal data, both running from −1 (perfect negative) through 0 (no linear relation) to +1 (perfect positive). Regression goes further and draws the line: simple linear regression fits Y = a + bX by least squares, so that b estimates how much Y changes per unit of X, and Y can be predicted from X. The coefficient of determination, r-squared, is the share of Y's variance explained by X — an r of 0.8 means 64% — and logistic regression extends the idea to binary outcomes such as disease yes/no, yielding adjusted odds ratios for each predictor. Correlation, however strong, never proves causation: that inference needs design, not arithmetic.

## What you must remember

- **Pearson's r:** parametric, for two continuous, roughly normal variables; sign gives direction, magnitude gives strength; r is unitless.
- **Spearman's rho:** rank-based non-parametric alternative for ordinal or skewed data and for outlier-laden series.
- **Interpretation bands:** 0 no linear relationship; below 0.3 weak, 0.3-0.7 moderate, above 0.7 strong — with the caveat that r detects only linear patterning.
- **r-squared:** variance explained; r = 0.8 → r² = 64%; the companion statistic to any correlation quoted in a paper.
- **Simple linear regression:** Y = a + bX; b (regression coefficient) is the change in Y per unit X; a is the intercept; fitted by least squares, minimising squared residuals.
- **Multiple regression:** several predictors of a continuous outcome; logistic regression for binary outcomes reports adjusted odds ratios, the workhorse of clinical research tables.
- **Correlation is not causation:** confounding, reverse causation and coincidence all generate strong r; only an experiment or rigorously designed cohort can argue causality.
- **Scatter first:** always inspect the scatter diagram — a strong curvilinear relationship (drug effect versus dose) can yield r near 0 and still be a real relationship.

## Worked example: birth weight and gestational age

Plot 200 newborns with gestational age on the x-axis and birth weight on the y-axis: the cloud rises from lower left to upper right. Pearson's r computes to 0.8 — strong positive — and r-squared says gestational age explains about 64% of birth-weight variation; the remaining 36% belongs to maternal nutrition, sex, parity and measurement noise. Now fit the regression: Y = −2,500 + 165X (weight in grams, age in weeks), meaning each extra week in utero adds about 165 g, on average, within the observed range — the phrase "within the observed range" matters, because extrapolating to 60 weeks is meaningless. If instead the outcome were low birth weight (yes/no) predicted from maternal haemoglobin, weight and smoking, the tool changes to multiple logistic regression, and each variable reports an adjusted odds ratio with its confidence interval, holding the others constant — which is precisely how confounders are tamed in analysis.

The teaching sequence in these two paragraphs is the one to reproduce in short answers: plot, quantify (r, r²), model (b, or adjusted OR), interpret within range, and never let the arithmetic outrun the design.

## How the exam frames it

Three recurring items. First, the value-of-r questions: r = −0.9 describes a strong inverse relationship, and candidates who miss the sign miss the mark; r = 0 with a U-shaped scatter tests whether you know r measures only linear association. Second, the correlation-versus-regression distinction: correlation is symmetric (r of X-with-Y equals Y-with-X) while regression is not — regressing Y on X yields a different line from X on Y — and prediction belongs to regression alone. Third, the causation trap: a vignette correlating television ownership with coronary deaths expects "confounding by affluence", not "televisions cause heart disease". The r-to-r-squared conversion (0.8 → 64%, 0.5 → 25%) is the cheapest mark in the chapter and is asked nearly every year in some form.

## Frequently asked questions

### What does a Pearson correlation coefficient of −0.85 indicate?

A strong inverse linear relationship between two continuous variables — as one rises, the other falls, with about 72% of variance shared (r² = 0.72).

### When should Spearman's rank correlation be used instead of Pearson's?

For ordinal data, non-normally distributed variables, or series with influential outliers, since ranking blunts the effect of extreme values.

### What does the regression coefficient b represent in Y = a + bX?

The average change in Y for a one-unit increase in X, holding the model's assumptions; in multiple regression each coefficient is adjusted for the other variables.

### Why can correlation never establish causation?

Because association may reflect confounding, reverse causation or chance, and only an experimental or well-designed longitudinal design can support a causal claim.

### Which regression handles a binary outcome such as disease present or absent?

Logistic regression, which models the log-odds of the outcome and expresses each predictor's effect as an adjusted odds ratio.
