Statistics for Comparison of Data Sets

On this page
  1. Direct answer
  2. What you must remember
  3. Consistency against productivity
  4. Shifts, scales and CV
  5. Frequently asked questions
  6. Related topics

Direct answer

Two distributions can share a mean and differ in everything else, so comparison runs on spread and consistency: variance σ² = Σxi²/n − x̄², standard deviation its positive square root, and the coefficient of variation CV = (σ/x̄) × 100 measuring relative spread — smaller CV, more consistent. Comparisons must respect units: raw SDs from different scales cannot sit side by side, and CV exists to normalise by the mean. Transformations behave predictably — adding a constant moves the mean, median and mode but leaves the SD untouched; multiplying by a scales the mean and SD by a, the variance by a². For moderately skewed data, mean − mode = 3(mean − median) flags the direction of skew.

What you must remember

  • Variance, both faces: σ² = Σ(xi − x̄)²/n = Σxi²/n − x̄² — the second form computes faster from raw data and its popularity in JEE Main is total.
  • CV rule: CV = (σ/x̄) × 100; smaller CV means more consistent, and consistency is a property of relative spread, not of size.
  • Change of origin: x → x + c leaves σ and variance unchanged, shifting mean, median and mode by c — the deviations never notice.
  • Change of scale: x → ax multiplies mean and SD by a and variance by a²; combined y = ax + c gives σy = |a|σx.
  • Averages' hierarchy: mean uses every value but yields to outliers; median survives skew and open-ended classes; mode answers "most frequent".
  • Empirical relation: mean − mode = 3(mean − median) for moderately skewed unimodal data; right skew drags the mean above the median.
  • Combined groups: the pooled mean weights group means by sizes; combined variance needs the between-group term too — a separate computation, not an average of variances.

Consistency against productivity

Two batsmen across a season: A averages 45 with SD 5; B averages 55 with SD 12. The CVs: 100 × 5/45 = 11.1% for A against 100 × 12/55 = 21.8% for B — A is markedly the more consistent, B the bigger but wilder scorer. Which to pick depends on the question: consistency questions answer A, run-volume questions answer B, and the CV exists to keep the two conversations separate. Now the transformation subtlety worth internalising: convert A's innings from runs to "runs above 40" — mean 5, SD still 5 — and the CV jumps to 100%, though nothing about the batting changed. Shift the thermometer instead: temperatures with mean 30, SD 4 in Celsius become mean 86, SD 7.2 in Fahrenheit, and the CV slides from 13.3% to 8.4% with no physical change at all. The lesson is real: CV is meaningful on ratio scales where zero means zero, and pure scaling leaves it invariant while shifting does not — a subtlety the syllabus implies but rarely states.

Shifts, scales and CV

JEE Main asks the origin-scale effects as certainties: adding 5 to every observation changes the mean and leaves the SD (the top distractor scales the SD too); multiplying by 3 triples the SD and multiplies variance by 9. The CV computation itself appears with raw data — compute mean, compute Σxi², produce variance, produce CV — and the consistency verdict is the answer, not the numbers. JEE Advanced prefers the comparison reasoning: choosing the better factory, batsman or machine as the question defines "better", where the trap is declaring the larger-mean candidate "more consistent" by conflating the two criteria. The pooled-group questions belong to the combined-variance chapter's machinery but the comparison logic lives here: two sections with equal means can differ entirely in spread, and the correct ranking tool is the CV whenever units differ between the candidates. The empirical relation supplies the skew questions: given mode and median, estimate the mean, or identify the skew direction from mean greater than median. One discipline ties the chapter together: state which measure the question demands — level, spread or relative spread — before touching any formula.

Frequently asked questions

Which measure decides which distribution is more consistent?

The coefficient of variation CV = (σ/x̄) × 100 — the smaller CV marks the more consistent data set.

How does adding a constant affect mean and standard deviation?

The mean shifts by the constant; the SD and variance are unchanged because deviations from the mean are untouched.

How does multiplying by a constant affect the spread?

The SD scales by |a| and the variance by a², with the mean scaling by a as well.

When is the median preferred over the mean?

For skewed data or open-ended classes, where outliers and missing upper limits drag the mean but leave the median stable.

What does the empirical relation mean − mode = 3(mean − median) say?

It links the three averages in moderately skewed unimodal data and indicates skew direction: mean above median signals right skew.

Practise this in the PrepElephant app

Question banks, previous-year questions, mock tests and revision tools — for Statistics for Comparison of Data Sets and JEE Mathematics. Free to start.

Get the free app WhatsApp