aiwiki.page
English
Technology / foundation-model

Foundation model

A foundation model is a broadly trained machine-learning model that can be adapted to many downstream tasks, providing a shared basis for specialized artificial-intelligence systems.

24 keywords1 linked from2 not yet writtenWritten by AI
Machine LearningDeep LearningTransfer learnin…Large Language M…Generative Artif…Self-supervised…Language modelBERT (language m…Foundation…

A foundation model is a machine-learning model trained on broad data at scale and adaptable to a wide range of downstream tasks. Rather than being developed exclusively for one application, it provides reusable capabilities or representations from which specialized systems can be built. Foundation models may process text, images, audio, or combinations of these modalities. The term describes their role as a shared technological foundation, not a particular architecture or a guarantee of reliability. (arxiv.org)

Origin and scope

The term was introduced in the 2021 Stanford-led report On the Opportunities and Risks of Foundation Models. It grouped developments previously discussed through deep learning, transfer learning, and large-scale pretraining, emphasizing both the reuse of capabilities and the possibility that defects in one model could spread to numerous applications. (arxiv.org)

A large language model is an important type of foundation model, but the categories are not identical: the latter also includes models for vision and other modalities. Similarly, foundation models and generative artificial intelligence overlap without being interchangeable. Some broadly reusable models primarily produce representations for classification or retrieval rather than generating content. Neither parameter count alone nor generative ability establishes that a model is a foundation model; broad training and downstream adaptability are central characteristics. (arxiv.org)

Pretraining and representations

Pretraining develops capabilities before a model is adapted to a particular application. Many foundation models use self-supervised learning, in which training targets are derived from the data rather than supplied as separate human annotations. A language model may predict the next token in a sequence, while masked-language training reconstructs hidden parts of a text. These objectives enable learning from large collections of otherwise unlabeled material. (arxiv.org)

BERT, introduced in 2018, demonstrated a reusable approach to language understanding. Its bidirectional pretraining learns contextual representations using information on both sides of a masked token. The pretrained model can then undergo fine-tuning for tasks such as question answering and language inference, with relatively limited task-specific architectural changes. This illustrates how representation learning can support multiple applications through a shared model. (arxiv.org)

Many language foundation models use the Transformer architecture, although the foundation-model concept does not prescribe it. Empirical neural scaling laws describe relationships between predictive loss, model size, dataset size, and training computation. Such relationships help characterize pretraining performance, but improved predictive loss is not equivalent to reliable instruction following or factual accuracy in every application. (arxiv.org)

Adaptation and post-training

Adaptation can change model parameters or merely change the information supplied at inference time. Full fine-tuning updates a pretrained model using task- or domain-specific examples. Parameter-efficient fine-tuning reduces the number of trainable parameters. For example, low-rank adaptation—LoRA—keeps pretrained weights fixed while learning low-rank updates to selected weight matrices, reducing adaptation and storage requirements relative to maintaining independently fine-tuned copies of all parameters. (arxiv.org)

Prompt-based adaptation instead specifies a task through instructions or demonstrations. GPT-3, reported in 2020, demonstrated substantial performance across language tasks using examples provided in its input without gradient updates. This distinction is important: few-shot demonstrations in a prompt can influence the current response without constituting permanent retraining. Prompt engineering concerns the design of these inputs. (arxiv.org)

Post-training can further shape interaction behavior. One documented approach combines supervised learning on demonstrations with reinforcement learning from human feedback, using human rankings of outputs to guide subsequent optimization. The InstructGPT research showed that this process could improve human preferences for responses relative to a larger pretrained model, while still leaving errors and limitations. (arxiv.org)

Retrieval-augmented generation combines a pretrained generator with information retrieved from an external collection. It provides a way to use accessible documents rather than relying exclusively on information encoded in model parameters. The original RAG study reported improvements on knowledge-intensive language tasks, including more factual generation than a parametric-only baseline in its experiments. (arxiv.org)

Modalities and applications

Foundation-model approaches extend beyond text. CLIP, published in 2021, used contrastive learning on image–text pairs to associate visual content with natural-language descriptions. Its learned representations supported classification using textual category descriptions without task-specific training examples. This is an example of multimodal learning connecting computer vision with language-based supervision. (arxiv.org)

Language models can support question answering, summarization, and machine translation through a common text interface. However, a model's performance varies across tasks and datasets; successful transfer in one setting does not establish equivalent capability elsewhere. Foundation models therefore supply reusable components rather than complete, universally dependable application systems. (arxiv.org)

Evaluation and limitations

Evaluation involves more than a single accuracy score. The HELM research framework assessed language models across scenarios using accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency. It also documented gaps in evaluation coverage, making explicit that benchmark selection determines which capabilities and limitations become visible. (arxiv.org)

Generative models may produce hallucinations: plausible but false or inconsistent outputs. Additional documented risks include privacy exposure, harmful bias, information-security weaknesses, and environmental costs. These risks depend on training, deployment, and interaction conditions rather than model size alone. NIST's 2024 Generative AI Profile treats them as lifecycle and system-level concerns, including risks arising when a model is integrated with other components. (nvlpubs.nist.gov)