---
module: 101-02
language: en
chapter: 101
title: "Quantitative Biostatistics for Biomedical and Clinical Science"
module_title: "Hypothesis testing, effect sizes, power, and multiplicity"
source_sha256: de1b4d957e07fc042e8104d4c0e36bda254ec2bb87dba76e2d87c7fe256e1e19
---
# Hypothesis testing, effect sizes, power, and multiplicity

## Tests and p values
### Asks if data are incompatible with a model
#### Null: no difference, no association, reference value
### Does not decide theory, importance, or bias
### p: chance of a result at least as incompatible
#### Assumes the null model and all assumptions
### Not the probability the null is true
#### Nor chance occurrence, nor replication success
### Small p: effect, trivial effect, violation, bias, selection

## Error rates and power
### Significance level: prespecified long-run threshold
### Type one error: rejecting a true null
### Type two error: not rejecting a false null
### Properties of procedures, not labels for one result
### Power depends on size, effect, variability, design
#### Plan for an effect worth detecting
#### Post hoc power restates the p value
#### Low prior plausibility: fewer significant findings true
### One-sided only if opposite effect equals no effect
#### Unexpected harm is rarely irrelevant
#### Specify direction before seeing outcomes

## Effect sizes and clinical importance
### Statistical significance is not clinical significance
### Risk difference: absolute change
### Risk ratio: proportional change
### Odds ratio looks more extreme when events are common
#### Approximates risk ratio when events are rare
### Mean difference keeps natural units
### Standardised difference trades meaning for comparability
### Minimal important difference: smallest worthwhile change
#### Varies with severity, burden, cost, harms
### Compare interval with benefit, trivial, harm zones
#### Significant but trivial, or nonsignificant yet important

## Matching analysis to data structure
### Paired analysis removes stable between-unit variation
#### Independent test on paired data wastes information
#### Paired test on unrelated data invents dependence
### Repeated measures need correlation over time
#### Treating all as independent gives false precision
### Rank-based tests are not assumption-free
#### Large samples: mean comparisons fairly robust
### Chi-square needs adequate expected counts
### Software p values are not self-validating
#### Normality pretest makes unstable two-stage inference

## Multiplicity
### Many outcomes, times, subgroups, models, looks
#### Chance of at least one false rejection rises
### Family-wise error: any false rejection
### False discovery rate: expected false proportion
### Bonferroni divides threshold by number of tests
#### Conservative when tests correlate
### Stepwise procedures improve power
### No adjustment rescues hidden unsuccessful tests

## Selective reporting and subgroups
### Choices made after results are known
#### Researcher degrees of freedom yield convincing p values
### Prespecification, registration, protocols, code
### Exploration valuable if labelled and retested
### Subgroup claims need a direct interaction test
#### Not significant here and nonsignificant there
### Credible if prespecified, few, plausible, replicated
### Arbitrary cut points create apparent thresholds

## Equivalence and non-inferiority
### Reverse the usual burden
### Equivalence: whole interval within both margins
### Non-inferiority: exclude unacceptable loss
#### Margin preserves justified comparator benefit
### Bias toward similarity can fake non-inferiority
#### Poor adherence, outcome misclassification
#### Examine intention-to-treat and per-protocol

## Reproducibility and publication bias
### Computational: same data and code, same result
### Replicability: new data, compatible evidence
### Generalisability: other populations and settings
### Not compressible into p below a threshold
### Publication bias favours positive, novel results
#### Small-study effects from selective reporting
#### Funnel asymmetry cannot diagnose one cause
### Pooling biased studies: precisely wrong

## Sequential evidence and interpretation
### Repeated looks inflate false positives
#### Group-sequential boundaries, alpha spending
### Early stopping for benefit exaggerates effects
### Safety: weigh severity and pattern
### State effect, interval, model, comparison
#### Then examine bias and relevance
### Convergent evidence changes understanding
