Statistics is the discipline concerned with collecting, organizing, analyzing, and interpreting data. It develops methods for describing observed patterns, drawing conclusions about populations or processes, and assessing uncertainty. Its scope extends from the design of investigations to the communication of results: statistical conclusions depend not only on calculations but also on how observations were obtained and which assumptions are justified. Statistics provides a foundation for empirical research and contributes to data science and machine learning. (online.stat.psu.edu)
Populations, samples, and data
A population is the entire collection of units relevant to an investigation; a sample is the subset observed. Units may be people, manufactured objects, organizations, or repeated measurements of a process. A population parameter describes a population characteristic, whereas a statistic is calculated from sample data. Thus, a sample mean can serve both as a description of observations and as an estimate of an unknown population mean. (online.stat.psu.edu)
Variables record characteristics of these units. Categorical variables distinguish groups, such as material types, while quantitative variables express numerical amounts. The distinction affects which summaries and analyses are meaningful. Data collection also requires operational definitions: a variable’s interpretation depends on what was measured, how it was measured, and which units were included. (online.stat.psu.edu)
Description and exploration
Descriptive statistics summarizes the data actually observed. Measures of location include the mean and median; measures of dispersion include the range, variance, and standard deviation. Frequencies and proportions describe categorical observations. No single summary fully characterizes a dataset: collections with similar averages can differ substantially in variability or distributional shape. (itl.nist.gov)
Exploratory data analysis examines structure before committing to a particular model. Through data visualization, including histograms, box plots, and scatter plots, analysts investigate asymmetry, unusual observations, relationships, and possible departures from assumptions. Exploration can generate questions and suggest models, but patterns discovered in the same observations used to evaluate them require care in subsequent inference. (itl.nist.gov)
Data collection and study design
Statistical reasoning begins before analysis. Probability sampling selects units through a specified random mechanism, enabling assessment of sampling uncertainty. Coverage problems, nonresponse, and selection bias can undermine generalization; increasing sample size does not by itself remove these problems. The population to which findings apply must therefore be distinguished from the population actually accessible to the study. (online.stat.psu.edu)
Experimental design addresses how conditions are assigned and comparisons constructed. Random sampling concerns selection into a study; random assignment concerns allocation to conditions within it. The former supports population generalization, while the latter helps support conclusions about causation. A randomized controlled trial uses random assignment to compare interventions. Observational studies instead record conditions without assigning them, making alternative explanations and confounding especially important. (online.stat.psu.edu)
Probability and statistical inference
Probability supplies mathematical tools for representing uncertainty. A statistical model describes observations through random variables and their probability distributions. Statistical inference uses observed data, together with a sampling design or model, to estimate unknown quantities, test specified claims, or predict unobserved outcomes. Its validity depends on the adequacy of those foundations. (online.stat.psu.edu)
In frequentist inference, procedures are evaluated through their behavior under repeated sampling. A confidence interval expresses uncertainty using a procedure with a specified coverage rate. Under the relevant assumptions, a 95% procedure produces intervals containing the fixed parameter in approximately 95% of repeated applications. This is not ordinarily a statement that the parameter has a 95% probability of lying inside the particular interval already calculated. (itl.nist.gov)
Bayesian inference combines a prior distribution with the data’s likelihood function, using Bayes’ theorem to obtain a posterior distribution. Posterior probabilities describe uncertainty conditional on the model, prior, and observed evidence. A credible interval contains a specified amount of posterior probability and therefore has a different interpretation from a frequentist confidence interval. (itl.nist.gov)
Testing and modeling
Statistical hypothesis testing evaluates observations against a specified hypothesis and associated assumptions. A p-value is the probability, under the null model, of a test statistic at least as extreme as the observed value. It does not give the probability that the hypothesis is true, nor does it measure an effect’s magnitude or practical importance. Failure to reject a hypothesis is not proof of its truth. (amstat.org)
Models also describe relationships among variables. Linear regression expresses a response through coefficients multiplying explanatory variables, plus an error term. Ordinary least squares estimates coefficients by minimizing squared residuals. “Linear” refers to the parameters, so the explanatory terms may include transformations or powers. Model assessment considers residual patterns, unusual observations, and the risks of extrapolating beyond the observed range. (itl.nist.gov)
Interpretation and applications
Statistics supports scientific investigation, industrial process control, surveys, and economic measurement. Its contribution to data science includes uncertainty quantification, sampling design, and causal inference, alongside computational methods for prediction. These activities overlap, but a useful prediction and a justified explanation are not interchangeable goals. (itl.nist.gov)
Statistical reporting distinguishes effect magnitude from uncertainty and states the assumptions and limitations underlying conclusions. Multiple analyses and selective reporting can distort the interpretation of apparently compelling results. Transparent descriptions of methods, data selection, and analytical decisions support reproducibility and allow readers to evaluate evidence beyond a single numerical threshold. (amstat.org)