ISEGORIA / MATH ENCYCLOPEDIA
Probability: learning from uncertainty
Bayesian updating, sums of random variables, and conditional probability.
Before you begin: Fractions, counting, and basic probability
Predict, manipulate, then check your reasoning against the example and question. Graphs illustrate the mathematics; they do not replace a proof.
1. Update a belief
Choose a beta prior and counts of heads and tails. The solid posterior is the prior multiplied by the likelihood and renormalized; the shaded band is its central 95% credible interval, and the gold line marks the posterior mean. Play the flips to watch the belief update one observation at a time.
Worked example. The uniform prior Beta(1, 1) with 3 heads and 1 tail gives Beta(4, 2), whose mean is 2/3.
Watch out. The model assumes independent flips with a fixed unknown probability. The posterior is conditional on these assumptions.
Why is the posterior mean not exactly the observed fraction?
The prior also contributes information; its relative influence decreases as observations accumulate.
2. Standardized sums
Increase the number of independent Bernoulli trials. On the left, bars show the exact binomial mass of the standardized sum, multiplied by σ so that bar areas are probabilities; the curve is the standard normal density. On the right, the largest gap between the two distribution functions is computed for every n and compared with the Berry–Esseen bound.
Worked example. For n=20 and p=0.5, the sum has mean 10 and variance 5.
Watch out. The central limit theorem concerns distributional convergence. Small n or highly skewed trials can give poor approximations.
Does the theorem say individual trials become normally distributed?
No. The standardized sum approaches a normal law; each trial remains Bernoulli.
3. Conditional probability
The unit square is the sample space and area is probability. Event B is a vertical strip; drag its edge, and drag the heights of A inside and outside B. On the right, the same intersection A∩B is divided by P(B) to give P(A | B) and by P(A) to give P(B | A).
Worked example. If P(B)=0.4 and P(A|B)=0.75, then P(A∩B)=0.3.
Watch out. Conditional probability is not generally symmetric: P(A|B) and P(B|A) can differ.
When are A and B independent here?
When the two conditional heights are equal, so knowing B does not change the probability of A.