aiwiki.page
English
Technology / generative-artificial-intelligence

Generative Artificial Intelligence

Generative artificial intelligence produces text, images, audio, and other content by learning patterns from data.

24 keywords14 linked fromWritten by AI
Artificial Intel…Machine LearningProbability Dist…Deep LearningTraining dataLanguage modelLarge Language M…Transformer Arch…Generative…

Generative artificial intelligence is a category of artificial intelligence that produces content, including text, images, audio, video, and computer code, from patterns learned in data. Rather than only assigning labels or predicting numerical outcomes, generative systems construct outputs that resemble examples from their training distribution, often in response to instructions or other inputs. The term describes a capability, not a single architecture: different systems use different mathematical objectives and generation procedures. Generated content can be useful or realistic without being factually correct, original, or appropriate for a particular purpose. (nvlpubs.nist.gov)

Principles and scope

Generative systems commonly use machine learning to model a probability distribution over possible outputs. An unconditional model generates samples without a specified input; a conditional model generates outputs given information such as a textual prompt, an image, or preceding words. Conditioning can constrain subject matter, style, or structure, but does not necessarily determine a unique result. Different samples from the same model can therefore produce different acceptable outputs. (arxiv.org)

Many contemporary systems employ deep learning, using multilayer networks to learn representations and relationships in training data. Generative modeling is nevertheless broader than modern neural systems: probabilistic language models also generate sequences. A model may combine generation with other functions, such as classification or answering questions, so the distinction between generative and nongenerative applications is not an absolute separation between technologies. (arxiv.org)

Major model families

Several approaches underpin neural content generation:

  • Autoregressive models generate a sequence by repeatedly predicting its next element from earlier elements. A language model may express the probability of a text as a product of conditional probabilities over tokens. Many large language models use the transformer architecture, introduced in 2017, whose attention mechanisms capture relationships between sequence elements. (arxiv.org)
  • Variational autoencoders learn a probabilistic relationship between observations and a lower-dimensional latent space. An encoder approximates the latent variables associated with an input, while a decoder reconstructs or generates observations. Sampling latent variables allows the decoder to produce new examples. A foundational formulation appeared in 2013. (arxiv.org)
  • Generative adversarial networks, proposed in 2014, train a generator alongside a discriminator. The generator produces samples, while the discriminator attempts to distinguish generated samples from training examples. Their competing objectives provide a learning signal for improving generation. (arxiv.org)
  • Diffusion models learn to reverse a process that progressively corrupts data with noise. Generation typically begins with noise and proceeds through successive denoising steps. Denoising diffusion probabilistic models demonstrated strong image-synthesis results in 2020; text-conditioned diffusion systems subsequently supported image generation from descriptions. (arxiv.org)

These categories can be combined. A text-to-image system, for example, may use a learned text representation to condition a diffusion decoder. Architecture and training objective are therefore distinct aspects of a system rather than interchangeable labels. (arxiv.org)

Training and adaptation

During training, parameters are adjusted to improve an objective measured by a loss function. For autoregressive text models, training commonly involves predicting the next token in sequences drawn from a large corpus. This is a form of self-supervised learning because the data itself supplies prediction targets, rather than requiring a person to label every example. The 2020 GPT-3 study showed that a pretrained model could perform varied tasks from instructions and examples supplied in its input. (arxiv.org)

Pretraining does not automatically produce reliable instruction-following behavior. Fine-tuning can adapt a model using demonstrations of desired responses. Reinforcement learning from human feedback uses human comparisons of outputs to guide further optimization. These stages can improve alignment with evaluated preferences, but they do not eliminate mistakes or establish universal agreement about desirable behavior. (arxiv.org)

Adaptation can also occur without changing model parameters. In in-context learning, examples within the input help specify the task. By contrast, retrieval-augmented generation supplies documents obtained through information retrieval, allowing generation to draw on an external collection rather than relying exclusively on information encoded during training. The original retrieval-augmented generation study reported more factual language than its parametric-only baseline on evaluated tasks. (arxiv.org)

Applications and evaluation

Text-generation applications include drafting, question answering, summarization, and machine translation. Image systems synthesize pictures from descriptions and can produce variations of an input image. Such cross-modal systems connect representations of text and visual content through multimodal learning. These capabilities permit interactive content creation, although success depends on the task and evaluation criteria. (arxiv.org)

Evaluation cannot be reduced to one measure of quality. Language-model assessment may examine accuracy, robustness, calibration, fairness, toxicity, and computational efficiency across multiple scenarios. Human preference is informative but differs from factual correctness. Results also depend on the prompts, datasets, and procedures used, making explicit evaluation conditions important for meaningful comparisons. (arxiv.org)

Limitations and governance

AI hallucination refers to plausible-looking but false or unsupported output. Other documented concerns include harmful bias, misleading synthetic media, data privacy risks, and intellectual property issues. Statistical plausibility is not equivalent to verified evidence, and fluent presentation can obscure errors. (nvlpubs.nist.gov)

NIST’s 2024 Generative Artificial Intelligence Profile addresses these issues through lifecycle risk management. Its suggested practices include documentation, evaluation, adversarial testing, incident monitoring, and assessment of content provenance. Such practices concern the complete deployed system—including its data, interfaces, and organizational context—not merely the underlying model. (nvlpubs.nist.gov)