---
module: 101-01
language: en
chapter: 101
title: "Quantitative Biostatistics for Biomedical and Clinical Science"
module_title: "Variables, distributions, sampling, and estimation"
source_sha256: 713180e1dd9679c97c77f784cf127a02fb4aafec62b80cfebde703ae1094fa6f
---
# Variables, distributions, sampling, and estimation

## Quantities before arithmetic
### Dilution conserves the solute amount
#### 2 mmol/L x 5 mL / 20 mL = 0.5 mmol/L
#### 20 mL is final volume, not solvent added
### Absolute versus relative change
#### 20 to 30: ten units, a fifty percent rise
#### 30 to 20: a one third fall, reference changes
### Percentages alone show neither reproducibility nor cause

## Probability-based reasoning and design
### From variable data to uncertain conclusions
#### Does not turn imperfect data into certainty
### Define population, sampling, variables, timing, target
### Correct maths can answer the wrong question
#### Eligibility, measurement, follow-up differ from target

## Variables and measurement
### Categorical: nominal or ordinal
#### Ordinal spacing not necessarily equal
### Quantitative: discrete counts or continuous
### Pain scale is ordered and analysed numerically
#### One unit may differ across range and people
### Accuracy, precision, reliability, validity
### Random error usually weakens associations
### Systematic error can create or conceal associations
### Calibration plus evidence for subjective constructs

## Distributions and summaries
### Centre: mean, median, mode
#### Mean uses every value, sensitive to extremes
### Spread: range, interquartile range, variance, SD
### Centre without spread hides heterogeneity
### Normal: symmetric, set by mean and SD
#### Mixtures, detection limits, subgroups distort shape
#### An approximation justified by purpose
### Log transformation symmetrises multiplicative variation
#### Back-transformed summaries read as ratios

## Z scores and reference intervals
### Standard deviations from the reference mean
### Not the probability of pathology
#### Unusual yet healthy, common yet dangerous
### Reference intervals hold a central proportion
#### Some healthy outside, some diseased inside

## Count distributions and probability
### Binomial: successes in fixed independent trials
#### Mean is trials times event probability
### Poisson: counts per time, area, or exposure
#### Mean equals variance in the simplest model
#### Overdispersion and clustering need extensions
### Independence differs from mutual exclusivity
### Bayes' rule updates a prior by likelihood
#### Post-test probability depends on pre-test probability

## Sampling and the standard error
### Random sampling supports generalisation
#### Convenience samples may differ systematically
### Random allocation supports causal comparison
#### Allocation does not make participants representative
#### Sampling alone does not remove confounding
### Standard error is SD of the sampling distribution
#### SD for individuals, SE for estimate uncertainty
#### Halving SE needs about four times the sample
### Central limit theorem: standardised means approach normal
#### Does not normalise data or remove bias
#### Clustered data carry less independent information

## Interval estimation and Bayesian inference
### 95% CI: procedure covers truth in 95% of repeats
#### Realised interval: parameter fixed, not moving
### Narrow interval around bias still misleads
#### Nonresponse, loss, misspecification, clustering unseen
### Crossing the null does not prove no effect
### Bayesian: prior plus likelihood gives posterior
#### Direct probability statements about parameters
#### Sensitivity analysis for contested priors

## Estimators, resampling, and sample size
### Judge by bias, variance, consistency, robustness
#### Modest bias can reduce prediction error
#### No universally best estimator
### Bootstrap resamples with replacement
#### Must preserve clustering and dependence
### Neither repairs selection bias
### Sample size links effect, variability, events, missingness
#### Define the minimally important effect beforehand
#### Inflating the assumed effect does not create power

## Descriptive analysis and the estimand
### Describe before complex modelling
### Investigate outliers, do not auto-delete
#### Data error, unusual valid patient, or incomplete model
#### Exclusion rules chosen after seeing impact invite bias
### Name the estimand
#### Population mean, risk difference, effect of intervention
#### Then ask if design and analysis identify it
