Random Variables and Expectation
On this page
Direct answer
A random variable attaches a number to every outcome of a random experiment, and expectation is its long-run average: for a discrete variable taking values x with probabilities p(x), E(X) = Σ x·p(x), where the probabilities are non-negative and sum to 1. Variance measures spread through Var(X) = E(X²) − (E(X))² — the expectation of the square minus the square of the expectation, in exactly that order. Linearity makes manipulation cheap: E(aX + b) = aE(X) + b, while Var(aX + b) = a²Var(X), so a shift moves the centre but never the spread. For independent variables, expectations multiply: E(XY) = E(X)·E(Y), which is the engine behind the indicator-variable method JEE Advanced favours.
What you must remember
- Distribution discipline: a probability table is valid only when every p(x) ≥ 0 and Σp(x) = 1; JEE numerical questions routinely plant an unknown constant k in the table and test this sum before anything else.
- Expectation formula: E(X) = Σ x·p(x); for a fair die E(X) = 3.5 — an average need not be a possible outcome. Linearity E(2X + 3) = 2E(X) + 3 holds always, independence not required.
- Variance and SD: Var(X) = E(X²) − [E(X)]² ≥ 0, standard deviation σ = √Var(X). For a Bernoulli(p) variable, E = p and Var = p(1 − p) ≤ 1/4.
- Shift and scale: Var(aX + b) = a²Var(X). Adding 5 changes nothing in variance; multiplying by 5 scales variance by 25 and SD by 5 — a perennial trap.
- Independence: E(XY) = E(X)·E(Y) and Var(X + Y) = Var(X) + Var(Y) need independence (uncorrelated suffices for the variance rule); assuming it silently is the most common error.
- Indicator method: for "expected number of..." questions, write X = I₁ + I₂ + ... with I taking value 1 when the event occurs, so E(I) = P(event). Expected fixed points in a random permutation of n letters = n × (1/n) = 1.
- Ordering facts: mean deviation about the mean never exceeds the standard deviation, and SD ≤ (maximum − minimum)/2 — quick elimination tools for multiple-correct questions.
A worked distribution table
Take a bag with 3 red and 2 black balls; two balls are drawn without replacement and X counts reds. The distribution is P(X = 0) = 1/10, P(X = 1) = 6/10, P(X = 2) = 3/10 — probabilities from combinations C(3, r)·C(2, 2 − r)/C(5, 2), already summing to 1. Then E(X) = 0(1/10) + 1(6/10) + 2(3/10) = 1.2, and E(X²) = 0 + 1(6/10) + 4(3/10) = 1.8, giving Var(X) = 1.8 − 1.44 = 0.36.
Notice the shortcut hiding in the arithmetic: without replacement, the mean is still n × (K/N) = 2 × (3/5) = 1.2 — the hypergeometric mean that mirrors the binomial np; only the variance changes. The habit worth building: compute E(X²) and E(X) separately, then subtract in the correct order — reversing the subtraction is how negative "variances" appear on answer scripts.
How JEE frames expectation
JEE Main stays at the table level: find k, find E(X), find Var(X) from an explicit distribution — pure plugin work under two minutes. JEE Advanced hides the distribution inside a game or a process: a coin tossed until a head appears, a player paid by the square of a die score, letters pushed into random envelopes. There the indicator method decides whether the attempt takes three minutes or fifteen. Two traps claim most marks. First, using E(XY) = E(X)E(Y) for dependent variables — the product rule is not a default. Second, answering a variance question with E(X²) alone: in numerical-answer format, that slip converts a correct value into a wrong one with no partial credit.
Frequently asked questions
How do you find an unknown constant in a random variable's probability distribution?
Impose Σp(x) = 1: if P(X = x) is proportional to x for x = 1 to n, then k·(1 + 2 + ... + n) = 1 gives k = 2/(n(n+1)), and every probability follows.
What is the fastest formula for the mean when sampling without replacement?
E(number of successes) = n × K/N, where n draws are made from K successes in N items — the exact analogue of np for binomial sampling.
Why is Var(aX + b) = a²Var(X) and not a²Var(X) + b?
Variance measures deviations from the mean; adding b shifts every value and the mean equally, so deviations are unchanged, while scaling by a multiplies every deviation by a and its square by a².
When does E(XY) = E(X)·E(Y) hold?
When X and Y are independent (uncorrelated is enough); for dependent variables the identity generally fails and must not be assumed.
What is the expected number of letters that reach the right envelopes in a random rearrangement?
Exactly 1, for any n: each of the n letters has probability 1/n of matching, and expectations add regardless of dependence.