Econometrics is the branch of economics that uses statistical methods, probability theory, and mathematical models to study economic data. It connects theoretical explanations of economic behavior with observed outcomes, estimating relationships, testing hypotheses, forecasting, and investigating causal effects. Its defining concerns include how economic variables are generated, which relationships can be identified from available evidence, and how uncertainty should be quantified. Econometric analysis encompasses both observational data and experiments; it is not restricted to regression or causal inference. (users.ssc.wisc.edu)
Historical development
A foundational development was Trygve Haavelmo’s The Probability Approach in Econometrics (1944), which formulated economic modeling within a probabilistic framework. Economic observations could thereby be analyzed as outcomes of stochastic mechanisms rather than as imperfect measurements of entirely deterministic relationships. Haavelmo received the Nobel Memorial Prize in Economic Sciences in 1989 for clarifying econometrics’ probability foundations and analyzing simultaneous economic structures. (arxiv.org)
Subsequent developments expanded methods for dynamic data, policy evaluation, and heterogeneous responses. The 2021 economic sciences prize recognized David Card’s empirical contributions to labor economics and Joshua Angrist and Guido Imbens’s methodological contributions to analyzing causal relationships. Their work clarified what researchers can infer from natural experiments, including situations where an intervention changes participation for only some individuals. (nobelprize.org)
Data and economic models
Econometric data commonly take three forms. Cross-sectional data describe different units, such as households or firms, at a particular period. Time-series data follow variables through successive periods. Panel data repeatedly observe the same units, allowing researchers to distinguish within-unit changes from differences between units. Each structure entails different assumptions about dependence and sampling. (users.ssc.wisc.edu)
A familiar model is linear regression:
Here, is an outcome, the are explanatory variables, the are coefficients, and collects influences not explicitly represented. Ordinary least squares estimates coefficients by minimizing squared residuals. Coefficients describe conditional associations unless further assumptions justify a causal interpretation. Econometric models can also be nonlinear, involve multiple equations, or describe discrete choices and dynamic adjustment. (users.ssc.wisc.edu)
Identification and causal interpretation
Identification asks whether a quantity of interest is uniquely determined by the distribution of observable data under stated assumptions. It differs from estimation: identification concerns what can, in principle, be learned, whereas estimation concerns extracting that information from a finite sample. Different causal explanations may generate the same observable associations. (arxiv.org)
A central difficulty is endogeneity, which arises when explanatory variables are correlated with a model’s disturbance. Sources include omitted common causes, certain forms of measurement error, and simultaneous determination. For example, observed prices and quantities reflect both supply and demand, so their correlation does not by itself identify either curve. Adding observations cannot automatically resolve this problem. (ocw.mit.edu)
Instrumental variables use an additional variable that shifts an endogenous explanatory variable while satisfying restrictions on its relationship with the outcome and unobserved determinants. Relevance alone is insufficient: the instrument must also support the required exclusion and exogeneity assumptions. With heterogeneous responses, an instrument may identify an effect for a particular subgroup rather than for the entire population. (arxiv.org)
Other approaches investigate causal relationships through research design. A randomized controlled trial assigns treatment randomly. Difference-in-differences compares changes in treated and comparison groups, typically requiring parallel untreated trends. Regression discontinuity exploits a treatment-assignment threshold, identifying effects near that threshold when appropriate continuity assumptions hold. These methods identify different quantities and rely on different assumptions; none is valid solely because a particular estimator is applied. (ocw.mit.edu)
Estimation and statistical inference
After specifying the target and identification assumptions, researchers choose an estimation procedure. Alongside least squares and instrumental-variable estimators, econometrics uses maximum likelihood estimation, generalized method of moments, and Bayesian inference. Their applicability depends on the model, available information, and assumptions about the data-generating process. (users.ssc.wisc.edu)
Statistical inference assesses uncertainty through standard errors, confidence intervals, and hypothesis tests. Economic observations may exhibit unequal disturbance variances, serial dependence, or correlation within groups. Robust or clustered standard errors address specified forms of this dependence, but they do not repair invalid causal identification. Sensitivity analysis examines how conclusions change under alternative specifications or assumptions. Statistical significance is distinct from economic importance: a precisely estimated effect can still be substantively small. (pubs.aeaweb.org)
Dynamics, forecasting, and machine learning
In macroeconomics and finance, time-series models address persistence, trends, and changing volatility. Regressing unrelated nonstationary series can produce apparently significant but misleading relationships. Cointegration describes circumstances in which a combination of nonstationary series is stationary, permitting analysis of long-run relationships alongside short-run adjustment. Conditional heteroskedasticity models represent variation in volatility over time. These contributions were recognized by the 2003 economic sciences prize. (nobelprize.org)
Machine learning contributes flexible prediction methods and tools for high-dimensional data. Regularization and cross-validation help manage model complexity and evaluate predictive performance. Econometric applications also use these tools to estimate intermediate components of causal models and investigate heterogeneous effects. Nevertheless, accurate prediction does not establish the consequences of intervention: a model can predict outcomes well while failing to identify how those outcomes would change under a different policy. (arxiv.org)