A generalized linear model (GLM) is a framework in statistics that extends linear regression to responses whose distributions need not be normal and whose means need not depend directly and linearly on predictors. It combines a response distribution, a linear predictor, and a link function connecting the predictor to the conditional mean. John Nelder and Robert Wedderburn introduced the unified framework in their 1972 paper “Generalized Linear Models,” bringing several established regression methods under a common estimation theory. (rss.onlinelibrary.wiley.com)
Model structure
A conventional GLM has three components:
Random component: conditional on the predictors, responses are independent and follow a specified probability distribution belonging to an exponential family.
Systematic component: the linear predictor is
Collectively, these predictors form , where is the design matrix.
Link component: a link function relates the conditional mean to the predictor:
The distribution and link are distinct choices, although particular combinations are customary. (arxiv.org)
“Linear” refers to linearity in the unknown coefficients, not necessarily in the original explanatory variables. Predictors can include powers, transformations, categorical indicators, and interaction terms. Thus, a model containing remains a GLM when its coefficient enters linearly. The link transforms the mean rather than the observed response: modeling is not equivalent to applying ordinary regression to . (search.r-project.org)
Distributional formulation
A common exponential-family representation is
where denotes a density or probability mass function, is the natural parameter, and is a dispersion parameter. The functions , , and specify the family. Differentiation gives
Consequently, the conditional variance is generally determined by the mean and dispersion rather than being constant. (arxiv.org)
A canonical link satisfies , so the natural parameter equals the linear predictor. Canonical links simplify estimation equations, but they are not compulsory; another valid link may provide a different relationship between predictors and the response mean. (maths.usyd.edu.au)
Important special cases
Several familiar models are GLMs:
| Response family | Common link | Conditional variance | Model |
|---|---|---|---|
| [[normal-distribution | Normal]] | Identity: | |
| [[bernoulli-distribution | Bernoulli]] | [[logit | Logit]]: |
| [[poisson-distribution | Poisson]] | Log: | |
| [[gamma-distribution | Gamma]] | Inverse or log |
The identity, logit, and log links are canonical for the normal, Bernoulli, and Poisson families, respectively. For Gamma responses, the inverse link is canonical, while the log link is also widely implemented. (maths.usyd.edu.au)
For grouped binary observations, a binomial distribution models successes out of a known number of trials; the link is conventionally applied to the success probability. A logit coefficient describes a change in log odds, and its exponential gives an odds ratio for a unit predictor change, holding other model terms fixed. Similarly, a log-link coefficient gives a multiplicative change in the conditional mean. These interpretations follow algebraically from the respective link equations. (search.r-project.org)
Estimation and inference
GLM coefficients are commonly fitted by maximum likelihood estimation. With independent responses, the joint likelihood is the product of individual response probabilities or densities. Gaussian regression with an identity link and equal variances yields the same coefficient estimates as ordinary least squares. Other families generally require numerical iteration. (maths.usyd.edu.au)
A standard method is iteratively reweighted least squares (IRLS). Each iteration constructs a working response and weights from current fitted means, solves a weighted least-squares problem, and updates the coefficients. This connects GLM estimation to familiar linear-regression computations without requiring normally distributed outcomes. Software implementations may encounter nonconvergence or fitted probabilities numerically close to zero or one. (rss.onlinelibrary.wiley.com)
Estimated dispersion contributes to coefficient standard errors. For ordinary binomial and Poisson families it is fixed at one, whereas Gaussian and Gamma models ordinarily estimate it. Pearson and deviance residuals provide alternative descriptions of discrepancies between observed responses and fitted means. (stat.ethz.ch)
Dispersion and extensions
Overdispersion occurs when variation exceeds that specified by the fitted response family—for example, when count variability exceeds a Poisson model’s mean–variance relationship. Quasi-likelihood methods specify a mean and variance relationship without requiring a complete response distribution; negative-binomial count models offer another approach. (arxiv.org)
A generalized linear mixed model adds random effects to accommodate clustered or repeated observations. Penalized GLMs introduce regularization, including ridge and lasso penalties, into likelihood-based fitting. These retain the distribution-and-link structure while modifying dependence assumptions or coefficient estimation. (arxiv.org)
References
- Generalized Linear Modelsmaths.usyd.edu.au
- R: Family Objects for Modelssearch.r-project.org
- R: Fitting Generalized Linear Modelssearch.r-project.org
- R: Summarizing Generalized Linear Model Fitsstat.ethz.ch
- R: Accessing Generalized Linear Model Fitsstat.ethz.ch
- A Family of Generalized Linear Models for Repeated Measures with Normal and Conjugate Random Effectsarxiv.org
- The family Argument for glmnetspout.uits.iu.edu