---
module: 101-03
language: en
chapter: 101
title: "Quantitative Biostatistics for Biomedical and Clinical Science"
module_title: "Regression, survival, prediction, causality, and missing data"
source_sha256: 9473452c9e296d3b0a9627fdb6a18492d9de58ba14730207a49faedb22396cdd
---
# Regression, survival, prediction, causality, and missing data

## Regression models and linear regression
### Describe, adjust, test interaction, predict
### Adjustment does not make a coefficient causal
### Scale, form, sampling, timing define parameters
### Linear: conditional mean of a continuous outcome
#### Coefficient: difference per unit, others fixed
### Check residuals: nonlinearity, variance, dependence
#### Residual normality matters mainly for small samples

## Predictor form and interaction
### Categorical predictors compare with a reference
### A straight line assumes constant per-unit effect
### Splines, polynomials, transforms model curvature
### Categorising loses information and power
#### Thresholds depend on chosen cut points
### Interaction: association varies by another predictor
#### Changes the effect on the model's own scale
#### Absent on one scale, possibly present on another

## Logistic and count models
### Logistic models the log odds of a binary outcome
#### Exponentiated coefficients are odds ratios
#### Common outcomes: relative change looks exaggerated
### Linearity assumed on the log-odds scale
### Poisson or negative binomial for counts
### Rates need numerator, person-time, recurrence rule
### Repeated events within a person are dependent
#### First-event analysis discards later burden
#### Cluster models: population-average or subject-specific

## Confounding and variable selection
### Common cause of exposure and outcome
### Adjustment blocks measured confounding paths
#### Adjusting a consequence removes part of the effect
#### Adjusting a common effect creates collider bias
### Causal diagrams make assumptions explicit
#### Identify an adjustment set but prove nothing
### Significance-based selection is unstable
### Choose confounders from causal knowledge
### Multicollinearity: imprecise separate coefficients
#### A signal of non-identifiability, not a disease

## Survival analysis
### Censoring: exact event time unknown
#### Assumed conditionally independent of events
#### Sicker participants leaving biases estimates
### Kaplan-Meier: event-free probability over time
#### Falls at events, not at censoring times
#### Median unestimable if curve stays above half
### Hazard: instantaneous rate among the event-free
### Cox model assumes a constant hazard ratio
#### Hazard ratio is not a risk ratio
#### Crossing hazards: restricted mean survival time
### Competing events are not always ordinary censoring
#### Censoring competitors estimates a hypothetical world
#### Cumulative incidence: observed-world probability

## Discrimination, calibration, and thresholds
### Discrimination: higher risk, more events
### Calibration: predicted versus observed risk
#### Good discrimination yet systematic overprediction
### Utility: do threshold decisions improve outcomes
### ROC curve ignores prevalence and error costs
#### Precision-recall helps with rare outcomes
### A generic index assumes equal error costs

## Overfitting and validation
### Model learns sample-specific noise
### Training performance is optimistic
### Resampling estimates optimism if whole process repeated
### External validation tests transportability
#### Case mix, measurement, treatment, prevalence differ
### Recalibrate intercept or slope
### Subgroup performance when harms are unequal

## Missing data
### Mechanisms defined relative to observed information
#### Completely at random: unrelated to any values
#### At random: depends only on observed variables
#### Not at random: depends on unseen values
### Assumptions about a process, not test results
### Complete-case analysis unbiased only narrowly
#### Loses precision, may shift target population
### Single imputation understates uncertainty
### Multiple imputation combines within and between

## Sensitivity analysis and clinical deployment
### Vary untestable assumptions plausibly
#### Missingness, unmeasured confounding, model form
### Robustness is not reporting only agreeing analyses
### Fragile if modest assumptions change the result
### Data leakage creates invalid performance
### Interpretability proves neither cause nor fairness
### Clinical use: validation, calibration, monitoring
### Align question, estimand, process, and model
