ISEGORIA / MATH ENCYCLOPEDIA
Statistics and inference: signals in data
Sampling, confidence, and regression.
Before you begin: Probability and algebra
Predict, manipulate, then check your reasoning against the example and question. Graphs illustrate the mathematics; they do not replace a proof.
1. Sampling means
Draw 200 samples of size n from a population and plot the histogram of their means. The means cluster around the population mean with spread σ/√n, whatever the shape of the population, and their histogram approaches the normal curve. Play to draw the samples one at a time.
Worked example. Larger samples reduce the standard error by a square-root law.
Watch out. Even a small standard error does not remove bias from a bad sampling frame.
What changes when n quadruples?
The standard error halves.
2. Confidence intervals
Fifty samples are drawn from a population with known σ, and each gives the interval x̄ ± zσ/√n. Move the confidence level: the samples stay the same, the critical value z changes, and every interval widens or narrows together. Count how many intervals miss the true mean.
Worked example. Higher confidence requires a wider interval under the same data model.
Watch out. The 95% describes the long-run coverage of the procedure over repeated samples; it is not a probability statement about this one fixed parameter.
What narrows the interval?
More observations or less variability.
3. Least-squares regression
Twelve points are drawn around the line y = 1 + slope·x with normal noise of the chosen scale. The least-squares line is refitted as you drag any point. Orange segments are residuals; the shaded band is the 95% confidence band for the mean response, and the right panel plots the residuals against x.
Worked example. The fitted line minimizes the sum of squared residuals.
Watch out. Association in a fitted line does not establish causation.
What does a residual measure?
Observed response minus fitted response.