High School Mathematics · Grades 10–12
34Sums of Random Variables and the Law of Large Numbers
Why do casinos always win in the end, and why do polls work? Because averages of many independent random quantities fluctuate less and less. This chapter proves it: linearity of expectation, additivity of variance for independent variables, the Bienaymé–Chebyshev inequality, and the law of large numbers.
34.1 Sums of random variables
Given two random variables on the same finite sample space , the sum is the random variable .
Theorem 34.1 (Linearity of expectation)
For all random variables on and :
No independence assumption is needed.
Definition 34.2 (Independent random variables)
and are independent if for all values :
Several variables are independent if this product rule holds for every choice of values of every subfamily.
Proposition 34.3
If and are independent, .
Proof.
∎
Theorem 34.4 (Variance of a sum)
If and are independent, then
More generally, for independent : .
Proof. Using König–Huygens (Proposition 33.3) and linearity:
and the last bracket vanishes for independent variables (Proposition 34.3). The general case follows by induction. ∎
Example 34.5 (Binomial revisited)
A binomial variable is a sum of independent Bernoulli variables. Hence, structurally:
recovering Theorem 33.8 without any computation.
Proposition 34.6 (Sample mean)
Let be independent with the same distribution as (expectation , variance ), and let be their sample mean. Then
Proof. Linearity gives ; independence gives , and dividing by scales the variance by (Proposition 33.4). ∎
The is the fundamental square root law: to halve the fluctuations of an average, quadruple the sample size.
34.2 Concentration inequalities
Theorem 34.7 (Markov’s inequality)
If and :
Proof. In (all terms nonnegative), keep only the terms with : each is at least , so . ∎
Theorem 34.8 (Bienaymé–Chebyshev inequality)
For any random variable and any :
Proof. Apply Markov’s inequality to the nonnegative variable with :
∎
Theorem 34.9 (Law of large numbers)
Let be independent with the same distribution (expectation , variance ) and their sample mean. For every :
The sample mean concentrates around the expectation.
Proof. Bienaymé–Chebyshev applied to , whose expectation is and whose variance is (Proposition 34.6). ∎
Remark 34.10
This theorem is the bridge between probability theory and statistics: the frequency of an event over many independent repetitions approaches its probability (take the indicator of the event, so that ). It justifies estimating a probability by simulation (Monte Carlo method) and a population proportion by a poll — quantitatively: see Chapter 35.
Method 34.11 (Using Bienaymé–Chebyshev)
To guarantee , it suffices that . The bound is crude (real fluctuations are usually much smaller) but perfectly general: it requires nothing about the distribution beyond a finite variance.
34.3 Exercises
Exercise 34.1 ★
Two fair dice are rolled; let be the sum. Using linearity (not the distribution of !), compute ; then compute using independence, given that a single fair die has variance .
Solution
Solution of Exercise 34.1.
Each die has expectation , so by linearity . The dice are independent, so .
Exercise 34.2 ★
Let . Bound using Markov’s inequality, then using Bienaymé–Chebyshev. Compare.
Exercise 34.3 ★
A fair coin is tossed times and denotes the frequency of heads. How large must be so that, by the Bienaymé–Chebyshev bound, ?
Solution
Solution of Exercise 34.3.
is the sample mean of Bernoulli variables, of variance . The bound
is as soon as .
Exercise 34.4 ★★
An investor splits her capital equally between independent assets, each with expected return and standard deviation . Compute the expectation and standard deviation of the portfolio return , and the number of assets needed to bring the standard deviation below . What financial principle does this illustrate?
Solution
Solution of Exercise 34.4.
By Proposition 34.6, (diversification does not change the expected return) and . Requiring gives , i.e. . This is the principle of diversification: spreading capital over independent risks divides the risk (standard deviation) by without reducing the expected return.
Exercise 34.5 ★★
Let and be independent, both uniform on .
- Give the distribution of and compute , directly from it.
- Recover both values by linearity and additivity of variance.
Solution
Solution of Exercise 34.5.
1. Counting the equally likely pairs: takes the values with probabilities . Hence and , so .
2. One variable: , , . Then and, by independence, . Same values.
Exercise 34.6 ★★
Show that the additivity of variance can fail without independence: compute for , and compare with . For which variables is ?
Solution
Solution of Exercise 34.6.
With : , while . The two agree only when , i.e. when is constant — so additivity genuinely requires independence (here is maximally dependent on itself).
Exercise 34.7 ★★
A die is suspected of being loaded. It is rolled times and shows a six times ( instead of ). Under the hypothesis that the die is fair, bound by Bienaymé–Chebyshev, and discuss whether the fairness hypothesis is plausible.
Solution
Solution of Exercise 34.7.
Under fairness, is the mean of Bernoulli variables with , :
The observed deviation is exactly : an event that a fair die produces with probability at most — and since Chebyshev is very conservative, the true probability is far smaller. The fairness hypothesis is not plausible; the die is very likely loaded.
Exercise 34.8 ★★★
(A better inequality for the coin.) Let and .
- Show that for .
- Deduce the distribution-free bound .
- How many people must be polled so that the observed frequency is within points of the true proportion with probability at least , using this bound? (Real polls use sharper estimates, but the order of magnitude is right.)
Solution
Solution of Exercise 34.8.
1. .
2. has expectation and variance ; Bienaymé–Chebyshev gives
valid whatever the unknown .
3. With and level :
About people suffice by this crude bound (the classical normal-approximation answer is nearer , see Chapter 35).
34.4 Problem: The house always wins
Problem 34.1
Weekend problem — the law of large numbers explains casinos, polls and insurance, and demolishes the gambler’s fallacy on the way
A roulette player after spins is, almost as often as not, ahead. The casino running a million spins is ahead with a certainty no court would question. Same game, same tiny edge — the difference is , and it is the subject of this chapter (Proposition 34.6, Theorem 34.8, Theorem 34.9). This problem runs the casino’s books, sizes an election poll, prices diversification — and dismantles the most expensive fallacy in the history of gambling.
Part I — Fluency.
- Two independent dice: compute (Theorem 34.4).
- Roll independent dice and average the results: give the expectation and the standard deviation of the sample mean (Proposition 34.6; for one die, ).
- A nonnegative random variable has expectation . What does Markov’s inequality (Theorem 34.7) say about ?
- Bound for the dice average of question 2 with Bienaymé–Chebyshev.
- Expectations add always, variances only under independence: exhibit dependent (hint: ) for which .
Part II — The house’s books. European roulette, bet euro on red: win with probability , lose with probability .
- Compute the expectation and standard deviation of one bet’s gain.
- A gambler makes independent bets; let be the total gain. Compute and , then the ratio . What does a drift of a quarter of a standard deviation mean for the gambler’s chances of being ahead tonight?
- The casino sees bets. Compute the expectation and standard deviation of its total take, and the same ratio. Interpret the number .
- Certify with Chebyshev: bound the probability that the casino loses money over the million bets.
- Betting systems: doubling after losses, quitting when ahead … Using the linearity of expectation (Theorem 34.1) over the (possibly random) sequence of stakes, explain why every strategy in a negative-edge game has negative expected gain — what would a winning system violate?
- State what the law of large numbers (Theorem 34.9) promises about the frequency of red — and what it does not promise about the next spin after ten reds in a row. Name the fallacy.
- The subtlest point: the law works by dilution, not compensation. If reds run ahead of expectation after some evening, the surplus is never “paid back” — compute what happens instead to the fraction by , and rewrite the gambler’s fallacy’s error in one sentence.
Part III — Polls.
- From Exercise 34.8: how many people must be polled for the observed frequency to fall within points of the truth with probability at least , by the distribution-free Chebyshev bound?
- Real polling institutes use about people for a -point margin at : their formula is with . Evaluate it, and explain the gap with question 13 (what does Chebyshev not know about the shape of the fluctuations?).
- To halve the margin of error, how must the sample grow? From points at , what sample gives points?
- The counterintuitive classic: the required sample size never used the population’s size — people suffice for a city or a continent. Point to the place in the model where the population size is absent, and give the kitchen analogy that pollsters use.
Part IV — Diversification.
- The insurer of Problem 33.1 earns euros in expectation on policies with a claims standard deviation of about . For which does the expected profit finally exceed one standard deviation of the claims? What is the doing for the insurer?
- One stock’s yearly return fluctuates with . Split the money equally over independent such stocks: compute the portfolio’s . Finance calls diversification the only free lunch — what is the lunch, exactly?
- Now let the stocks be perfectly correlated (they all move together): what is the portfolio’s ? Compare the two extremes and state which assumption every diversification argument secretly rents — and what happened when it failed system-wide in 2008.
- Finale — the symphony: sums fluctuate like while their means settle like ; play the theme through the four industries of this problem (casino, polling, insurance, portfolios), name the two standing abusers (the gambler’s fallacy and fake independence), and give the forward pointer: the shape of the fluctuations — the bell — is the next chapter’s continuous star, and its full theorem crowns the university volumes.
Solution
Solution of Problem 34.1.
1. (independence).
2. ; : the hundred-dice average hugs within a fifth of a point.
3. .
4. : at most about .
5. With : , while : perfectly anti-correlated variables cancel — variance addition is an independence privilege.
6. ; .
7. , : the drift is only . The night’s noise dwarfs the edge — a large minority of gamblers (roughly , says the bell) walk out ahead, which is precisely what keeps them coming back.
8. Casino’s take: expectation euros, standard deviation : the profit sits standard deviations above zero. At , “the casino might lose this year” is not a risk, it is a rounding error.
9. for the gamblers : even the bluntest inequality in the book guarantees the house at — the truth is astronomically stronger.
10. Each euro staked, whenever and however chosen, has expected return of itself; by linearity the expected total gain is , negative for every strategy that stakes anything. A winning system would violate the linearity of expectation — no clever sequencing of bad bets makes a good one. (Doubling systems merely trade many small wins for rare catastrophic losses.)
11. The theorem: the frequency of red over spins converges (in probability) to . It says nothing about spin : the wheel has no memory, and after ten reds the chance of red is still . Believing otherwise is the gambler’s fallacy, and question 12 shows what actually happens to streaks.
12. The surplus of reds is not repaid — future spins are fair copies, expected surplus stays . But : one hundredth of a point. The fallacy’s error in one sentence: the law of large numbers dilutes past accidents in an ocean of new trials; it never sends the wheel to collect debts.
13. people.
14. : five times fewer. Chebyshev holds for every distribution and pays for its universality with slack; the pollsters’ constant comes from the actual bell shape of the fluctuations, which concentrates far harder.
15. Margin : halving it quadruples the sample — about people for points. Precision is bought at quadratic prices.
16. The model is independent draws with success probability — the population size never appears (sampling a tiny fraction of a large, well-mixed population is what independence encodes). The pollsters’ analogy: to taste the soup, one well-stirred spoonful suffices — whether the pot serves ten or ten thousand. The catch is the stirring — a biased sampling frame, as in the Middle School volume’s poll that fooled a country — not the pot.
17. : beyond a hundred policies the drift outruns the noise, and each further policy widens the gap by the law — pooling is the business model.
18. Independent equal shares: : same expected return, one-fifth the fluctuation. The lunch: risk reduction at zero cost in expectation — averaging independent randomness.
19. Perfectly correlated: the portfolio is one stock in costumes: , no reduction at all. Diversification rents independence; when a crisis correlates everything (2008: all housing bets were one bet), the protection evaporates exactly when it is needed.
20. Sums drift like and fluctuate like , so means settle like : the casino banks the drift over a million spins; the pollster buys of noise for a phone budget; the insurer outgrows its own claims at ; the investor divides risk by . Abusers: the gambler who believes in debts (there is only dilution), and the financier who believes in independence (there is sometimes only one bet). The fluctuations’ universal shape — the bell — is the continuous chapter’s star, and its theorem, the central limit theorem, is the summit of the university volumes’ probability.