Calibration is an operation in metrology that establishes how the indications of a measuring instrument or system relate to values supplied by measurement standards, with measurement uncertainty taken into account. Under specified conditions, these relationships are used to obtain measurement results from subsequent indications. Calibration may produce a correction, table, curve, or mathematical relationship; it does not necessarily involve changing the instrument. The term also has a distinct use in statistics and machine learning, where it describes agreement between predicted probabilities and observed outcomes. (jcgm.bipm.org)
Meaning and related operations
The International Vocabulary of Metrology describes calibration as a two-stage operation. First, reference quantity values and their uncertainties are related to corresponding instrument indications and their uncertainties. Second, that information establishes how an indication can yield a measurement result. In everyday laboratory usage, “calibration” sometimes refers only to the comparison stage. (jcgm.bipm.org)
Calibration differs from adjustment, which changes a measuring system so that it gives prescribed indications. Altering a zero setting or sensitivity is an adjustment, whereas determining the correction needed for an unchanged instrument is calibration. Verification determines whether specified requirements are met, such as whether errors fall within permissible limits. A calibration result can support verification, but calibration itself is not simply a pass-or-fail judgment. (jcgm.bipm.org)
Reference standards and traceability
A reference standard supplies a quantity value with an associated uncertainty, rather than an absolutely error-free value. Metrological traceability connects a measurement result to a specified reference through a documented, unbroken chain of calibrations, each contributing uncertainty. References commonly include realizations of units in the International System of Units, although other explicitly defined references are possible. Traceability belongs to the measurement result, not merely to an instrument or its manufacturer. (nist.gov)
National metrology institutes maintain reference standards and provide calibration services that support this chain. Nevertheless, having an instrument calibrated by such an institute does not automatically establish traceability for every later measurement. The user’s measurement procedure, conditions, uncertainty evaluation, and continuing control of the equipment also matter. Nor does traceability alone establish that uncertainty is sufficiently small for a particular application. (nist.gov)
Calibration procedures
A procedure defines the measurand, or quantity intended to be measured, together with the measurement range, reference standards, comparison method, and relevant operating conditions. Calibration procedures cover quantities including mass, temperature, pressure, time, volume, and electrical quantities. Their technical details vary because different instruments and quantities require different comparison arrangements. (nist.gov)
Documentation records the method, standards, results, and uncertainty evaluation needed to support the claimed traceability. Comparisons at several reference values can establish the instrument’s response across a range. Repeated observations and continuing measurement-assurance checks help characterize variability and changes over time. Conditions and limitations are important: a calibration relationship established under one set of conditions is not an unrestricted guarantee of performance under every other condition. (nvlpubs.nist.gov)
Calibration curves and mathematical models
A calibration curve represents the relationship between reference values and instrument responses. A common model is
[ y=a+bx+\varepsilon, ]
where (x) is a reference value, (y) is the response, (a) is an offset, (b) is the response slope, and (\varepsilon) represents measurement error. This is a linear regression model; quadratic, power-law, and other nonlinear relationships are also used. Model selection cannot be justified by visual inspection alone. (itl.nist.gov)
For a later response (y'), the linear relationship gives the calibrated estimate
[ \hat{x}=\frac{y'-\hat{a}}{\hat{b}}, ]
provided the estimated slope is nonzero. Thus, applying calibration commonly involves the inverse of the response relationship. A nonlinear model may require choosing a physically appropriate root or using interpolation. The relationship can remain useful while the measurement process stays under statistical control; some chemical measurement procedures instead establish a fresh curve for each batch. (itl.nist.gov)
Uncertainty of calibrated results
Correcting an indication does not eliminate uncertainty. Uncertainty evaluation accounts for the reference values, observed variability, fitted relationship, and other relevant contributions. The uncertainty of a future corrected result is therefore distinct from the scatter of the original calibration observations alone. (itl.nist.gov)
The Guide to the Expression of Uncertainty in Measurement framework distinguishes Type A evaluations, based on statistical analysis of observations, from Type B evaluations, based on other information. Contributions are expressed as standard uncertainties and combined using a measurement model, including relevant covariances. These categories describe evaluation methods, not a simple division between random and systematic effects. (emtoolbox.nist.gov)
Expanded uncertainty is commonly expressed as (U=ku_c), where (u_c) is combined standard uncertainty and (k) is a coverage factor. Under suitable normal-distribution assumptions, (k=2) gives approximately 95 percent coverage. That interpretation is conditional, rather than universal. (emtoolbox.nist.gov)
Recalibration and continuing control
Calibration establishes a relationship at the time of measurement; instrument behavior can subsequently change. Recalibration intervals depend on stability, environmental influences, accuracy requirements, and applicable external requirements. There is no universally appropriate annual interval. Historical results, comparisons with other standards, and statistical control charts can provide evidence for setting or revising intervals. Records of performance before and after calibration help distinguish deterioration from changes introduced during servicing. (nist.gov)
Probability calibration
In predictive modeling, calibration concerns whether a predicted probability corresponds to observed frequency. For example, predictions assigned approximately 0.8 confidence should be correct approximately 80 percent of the time in the relevant group. Classification accuracy and calibration are distinct: a model can classify many cases correctly while expressing excessive confidence. Reliability diagrams compare confidence with empirical accuracy. (proceedings.mlr.press)
Post-processing methods fit a probability transformation using a separate validation set. Temperature scaling divides network logits by a learned positive scalar before applying the softmax function. This changes confidence without changing the highest-scoring class. Unlike physical calibration, this procedure relies on labeled observations rather than a traceability chain to measurement standards. (proceedings.mlr.press)