Permutations · Combinations · Binomial theorem
An interactive introduction to probability and statistics
Chance,
Data, and Decisions
Start with what might happen. Observe what did. Then learn how evidence turns uncertainty into a calibrated decision.
model
→simulate
→observe
→infer
→check
The course
One uncertainty unlocks the next.
The sequence follows a single intellectual arc: outcomes become variables, variables form distributions, distributions generate samples, and samples support inference.
Events · Birthday paradox · Derangements
A measure of possibility
Build a probability space from outcomes and weights, then count shared birthdays and misplaced hats.Conditioning · Independence · Bayes
Information changes the space
Restrict, renormalize, and reverse the direction of evidence.Discrete variables · Expectation · Variance
Numbers attached to outcomes
Turn outcomes into numbers, average their weights, and explain the spread of dice, payoffs, and counts.Uniform · Bernoulli · Binomial · Sampling
A library of chance mechanisms
Build discrete distributions from experiments, then move to waiting times and later continuous models.Joint laws · Marginals · Correlation
Variables together
See dependence as a whole surface, not a single number.LLN · CLT · Sampling distributions
From repetition to regularity
Watch averages stabilize and bell curves emerge.Estimation · Intervals · Tests
Learning from a sample
Turn random data into calibrated claims.Prior · Likelihood · Posterior
Probability for unknowns
Update uncertainty and predict what comes next.de Finetti · Pólya’s urn · Random limiting frequencies
Exchangeability: when symmetry produces a parameter
An optional capstone: connect symmetry, conditional independence, and learning through worked examples, nine exercises, and a birthday-collision proof.Interactive companion
Make randomness visible.
Every lab pairs a manipulable model with an interpretation question. Simulation is a microscope for assumptions—not a substitute for them.
Event lab
Shade the statement
Switch between words, symbols, and regions. Probability begins with getting the event right.
Read the region
A ∩ B: both events occurBefore computing, ask: are the outcomes ordered? Are they equally likely? Is the event inclusive or exclusive?
Counting lab
Does the order matter?
The same objects can produce very different sample spaces. Decide what makes two outcomes different.
Unordered · no repetition
56C(8, 3) = 8! / (3! 5!)A committee: swapping members changes nothing.
Ordered · no repetition
336(8)3 = 8! / 5!Distinct offices: changing who fills each role changes the outcome.
Ordered · repetition allowed
51283A sequence of draws with replacement.
Every 3-element subset has 3! = 6 orderings, so 56 × 6 = 336.
Read the binomial coefficient in Pascal’s triangle
Select an entry to change n and k. Each interior entry is the sum of the two above it: decide whether one distinguished object belongs to the subset.
C(8, 3) = C(7, 2) + C(7, 3) = 21 + 35 = 56
The binomial theorem counts the choices of factors that supply y: (x + y)n = Σk=0n C(n, k)xn−kyk. In particular, a row sums to 2n, the number of subsets.
Inclusion–exclusion lab
Count each outcome once.
Build three events from disjoint regions, then see the same overlap correction in counts and probabilities.
Eight disjoint regions · edit their numbers of outcomes
|A ∪ B ∪ C| = 65
Why add the triple overlap back? An outcome in all three events contributes 3 − 3 + 1 = 1. Every pairwise intersection includes the triple region.
The probability space behind the picture
Choose uniformly among the 100 individual outcomes. Each has probability 1/100; a region receives its count divided by 100. The eight regions can also be regarded as eight outcomes with those generally unequal masses.
Probabilistic inclusion–exclusion holds for arbitrary event probabilities, including nonuniform models. It does not require independence. In general, add singles, subtract pairs, add triples, and continue with alternating signs.
Birthday lab
How many people make a match?
A match somewhere in the room and a match with one particular person are different events.
Model: labeled people, independent birthdays, and 365 equally likely days. The 365-day version ignores leap days and seasonal variation.
At least one shared birthday
50.73%Count the complement: all birthdays distinct
P(some match) = 1 − (d)n / dn = 1 − ∏j=0n−1(1 − j/d)
Here (d)n = d(d − 1)⋯(d − n + 1) counts ordered selections without replacement. There are dn possible birthday lists in all. For one designated person, the answer is 1 − (1 − 1/d)n−1.
Predict the result, then run the experiment.
Three people: see inclusion–exclusion at work
Let E₁₂, E₁₃, E₂₃ say that the indicated pair matches. Each has probability 1/d. Any two force all three birthdays to agree, so every pairwise intersection and the triple intersection have probability 1/d².
P(a match among 3) = 3/d − 3/d² + 1/d² = 3/d − 2/d²
For this 365-day calendar, that is 0.82%. Multiplying the three no-match probabilities would assume mutual independence that these events do not have.
Bayes lab
Reverse the condition
A positive result is evidence, not a verdict. Watch the base rate reshape the answer.
Among positive results
8.8%actually have the conditionDistribution lab
Shape a probability law
Change the mechanism, then read its center and spread before looking at the graph.
Count successes in a fixed number of independent trials.
Dependence lab
Turn the correlation dial
Correlation captures linear alignment—not every form of dependence, and never causation by itself.
ρ
0.70moderate linear associationZero correlation only means no linear trend. A curve can still be perfectly predictable.
Sampling lab
Repeat the sample, not the slogan
See the CLT reshape averages, then watch confidence intervals succeed and occasionally miss.
2,600 simulated sample means
28 nominal 95% intervals · 27 cover μ
Bayesian lab
Watch evidence move a prior
Treat uncertainty about a proportion as a distribution, update it, then predict the next trial.
Posterior predictive chance of success on the next trial: 71.4%
Exchangeability capstone
Same marginal. Different long runs.
Every trial has success probability ½. Change how the trials share information and watch the count distribution change.
Begin with one red and one blue ball. Replace each drawn ball and add another of its color. This has the same word law as choosing one uniform bias on [0, 1] for the entire run.
Number of successes Sn · bar height uses the same 0–100% scale in every model · selected count 4
Next success, given this count
50.00%Order adds no information beyond the count in these exchangeable models.
Variance of Sn/n
0.10417Pairwise covariance: 0.08333.
Limit within a run
Uniform [0, 1]The limit is uniform on [0, 1]. Increasing n reduces sampling noise within the run.
Words and counts are different events. One specified word with 4 successes in 8 trials has probability 0.16%. There are 70 such words, so the probability of the count S8 = 4 is 11.11%. For the one-red, one-blue urn, all n + 1 counts are equally likely.
Read the probabilities and the variance identity
Var(Sn/n) = E[Θ(1 − Θ)]/n + Var(Θ).
The first term vanishes as n grows. The second describes variation of the shared bias between runs. In an infinite Bernoulli mixture, Cov(Xi, Xj) = Var(Θ) ≥ 0 for i ≠ j. This also explains why a finite always-disagreeing pair cannot extend to an infinite exchangeable sequence.
| Successes | Probability |
|---|---|
| 0 | 11.11% |
| 1 | 11.11% |
| 2 | 11.11% |
| 3 | 11.11% |
| 4 | 11.11% |
| 5 | 11.11% |
| 6 | 11.11% |
| 7 | 11.11% |
| 8 | 11.11% |
Continue in Sage. Download the exchangeability notebook for exact beta integrals, posterior predictions, seeded urn simulations, finite counterexamples, and a total-variation comparison bounded by birthday collisions. Select the SageMath kernel. Chapter 10 develops the proofs →
The foundations, worked through
From counting to a probability model.
Start with coins, dice, birthdays, and hats. Each chapter explains the model, works the calculation, and gives you problems to try.
Decide what you are counting.
Permutations, combinations, repeated objects, and selections with repetition. The binomial theorem follows by choosing which factors contribute each term.
Read the counting chapter →Give outcomes their weights.
An event is a collection of outcomes. Count equally likely outcomes or add their weights, then use complements and inclusion–exclusion for birthdays and derangements.
Read the probability models →Ask a numerical question.
A random variable records a number for each outcome. The event X ≥ a contains the outcomes whose numbers meet that threshold. Probability tables and CDFs describe its discrete distribution; weighted sums give its mean and variance.
Read the random-variable chapter →Chapter 4 · The next step
Find the average. Explain the spread.Expectation as a weighted sum, linearity, indicators, variance, and independent sums—all through discrete experiments.Work with dice, payoffs, correct hats, and matching birthdays, then try 18 discrete exercises with worked answers.Chapter 5 · What comes next
Give familiar experiments a distribution.Start with uniform choices, Bernoulli indicators, binomial counts, and sampling without replacement. Each example derives probabilities, a mean, and a variance.Geometric waiting times and Poisson counts follow; continuous distributions are marked for later reading.Chapter 5 · Coupon collection
How long until you have them all?Collect every coupon type—or roll every face of a die. Add geometric waiting times to find the mean and variance, then bound the chance of finishing by a deadline.Four exercises with worked answers and a bonus connecting collection to Poisson counts.Appendix A · Calculus refresher
Bring the calculus back into focus.Limits, derivatives, optimization, integrals, series, and Taylor error bounds, with 18 exercises and worked answers.Derive the birthday approximation and the derangement limit 1/e; then practice the calculus behind densities, means, and variances.Ready to work through a problem?
Browse student problem sets & quizzesPractice architecture
Calculate. Explain. Critique.
A complete assignment mixes technical fluency, mathematical reasoning, and model criticism.
Translate precisely
Move among words, events, tables, distributions, and formulas.
Example promptChoose a committee, assign officers, then explain the difference in the counts.
Justify the bridge
Derive results and identify exactly where assumptions enter.
Example promptWhy does the triple intersection have to be added back in inclusion–exclusion?
Stress-test a model
Simulate, visualize, compare, and diagnose a mismatch.
Example promptWhy is a birthday match anywhere in the room more likely than a match with one particular person?
Problem Set 2
Counting, events, and shared birthdaysProbability spaces · binomial coefficients · both forms of inclusion–exclusionProblem Set 3
Counting, collisions, and derangementsDiscrete probability spaces · poker and word counts · birthday bounds · derangementsComputational companion
Four Python notebooks and a Sage capstone
Labs 1–4 use Python. For Lab 5, choose the SageMath kernel and run from the top; the script also runs with sage filename.sage.
The governing question
What would we expect to see if our model were true?
That question joins probability, simulation, estimation, testing, and posterior prediction. It also keeps inference honest.