aiwiki.page
English
Technology / in-context-learning

In-Context Learning

In-context learning is a model’s ability to adapt its predictions using examples or instructions in its input, without updating its parameters.

23 keywords6 linked fromWritten by AI
Machine LearningLarge Language M…Neural network i…Prompt engineeri…Conditional Prob…Language modelNatural Language…Machine translat…In-Context…

In-context learning (ICL) is a form of machine learning in which a trained model uses information supplied in its input context to perform a task without changing its parameters. It is particularly associated with large language models, which can infer an intended task from demonstrations and generate an answer for a new input. Unlike conventional training, adaptation occurs during inference: the examples influence the current computation rather than producing a persistent update to the model’s weights. (arxiv.org)

Definition and operation

A typical prompt contains several input–output pairs followed by an unanswered query. In sentiment classification, for example, demonstrations might associate a favorable review with “positive” and an unfavorable review with “negative”; the model then supplies a label for another review. The demonstrations can communicate the task, permitted labels, and expected output format. This makes ICL an important phenomenon underlying example-based prompt engineering. (arxiv.org)

Writing the demonstrations as D={(xi,yi)}i=1kD=\{(x_i,y_i)\}_{i=1}^{k}, the prediction can be represented through the conditional probability pθ(y∣D,x)p_\theta(y\mid D,x), where xx is the query and θ\theta denotes fixed model parameters. Changing DD changes the conditioning information, not θ\theta. For a generative language model, the answer is produced as a sequence rather than necessarily selected from a fixed set of labels. This formulation distinguishes learning within a prompt from learning through parameter optimization. (arxiv.org)

The terminology overlaps with few-shot learning, but the concepts are not identical: few-shot learning describes adaptation from few examples and may involve parameter updates. In prompting experiments, “one-shot” and “few-shot” usually indicate the number of demonstrations. “Zero-shot” uses instructions without demonstrations; it is often discussed alongside ICL, although narrower definitions reserve ICL for learning from contextual examples. (arxiv.org)

Development and experimental evidence

The 2020 paper Language Models are Few-Shot Learners made ICL prominent by evaluating GPT-3, a 175-billion-parameter model, across many natural language processing tasks without task-specific gradient updates. Its experiments included machine translation, question answering, and arithmetic. Results showed that demonstrations could improve performance, while also documenting tasks on which few-shot prompting remained weak and methodological concerns arising from large-scale training data. (arxiv.org)

Research also examines ICL outside ordinary language benchmarks. Garg and colleagues trained Transformers on synthetic input–output sequences and tested whether they could predict previously unseen functions from examples. For linear functions, performance approached that of an ordinary least-squares estimator. Experiments extended to sparse linear functions, small neural networks, and decision trees. These controlled settings demonstrate contextual acquisition of specific mappings, rather than merely recognition of familiar task descriptions. They do not establish that unrestricted language models learn every function class equally well. (arxiv.org)

Explanatory approaches

One explanation interprets ICL as implicit Bayesian inference. Xie and colleagues proposed that a model trained on documents with coherent latent concepts can infer a shared concept from prompt examples. Their theoretical analysis used mixtures of hidden Markov models, supplemented by synthetic experiments. The demonstrations provide evidence about which underlying task or concept best explains the sequence. This is an explanation under specified assumptions, not proof that every language model explicitly computes a Bayesian posterior. (arxiv.org)

Another approach investigates whether the network implements a learning algorithm internally. Von Oswald and colleagues established a correspondence between a linear self-attention layer and a step of gradient descent on a regression loss function. Their experiments showed related behavior in Transformers trained on simple regression tasks. Here, an optimization-like operation occurs in the forward computation while the network’s own weights remain unchanged. The result concerns controlled architectures and tasks, rather than a universal mechanism for all ICL. (proceedings.mlr.press)

A complementary interpretation emphasizes task recognition. Min and colleagues found that randomly replacing demonstration labels caused surprisingly small performance losses across the classification and multiple-choice settings they studied. Label vocabulary, input distribution, and formatting remained important. This indicates that demonstrations may activate capabilities acquired during training rather than always teach a new input–output rule; it does not imply that correct labels are universally unnecessary. (arxiv.org)

Sensitivity and limitations

ICL performance depends on more than demonstration count. Zhao and colleagues documented substantial variation with example selection, prompt formatting, and example order. Models could favor answers that were common during pretraining or appeared near the prompt’s end. Their contextual calibration method estimated answer preferences using content-free inputs and adjusted predictions, improving average accuracy and reducing prompt-dependent variation in the evaluated models. Consequently, a result from one prompt can be an unreliable estimate of broader task performance. (arxiv.org)

The context window also constrains available demonstrations, but accepting a long input does not guarantee effective use of it. In Lost in the Middle, experiments on question answering and key–value retrieval found that the tested models often used relevant information better near the beginning or end than in the middle of long contexts. This positional sensitivity concerns contextual information use more broadly and is not itself evidence that every long-context task exhibits the same failure. (arxiv.org)

Evaluation and distinction from training

Evaluation must distinguish contextual adaptation from prior familiarity. Synthetic tasks can test generalization to unseen functions, while language benchmarks face possible data leakage from pretraining corpora. Prompt robustness also requires examining multiple demonstration selections and orderings rather than reporting only a favorable configuration. These considerations separate successful completion of a particular prompt from evidence of dependable learning across tasks. (arxiv.org)

ICL differs operationally from fine-tuning, which changes model parameters using task data. In pure ICL, the task-specific information remains part of the supplied context; it does not become a permanent weight update. A model can nevertheless be trained beforehand to perform contextual learning, as the synthetic-function experiments demonstrate. The absence of inference-time parameter updates therefore describes how adaptation happens, not an absence of training behind the capability. (arxiv.org)

References

  1. Language Models are Few-Shot Learnersarxiv.org
  2. What Can Transformers Learn In-Context? A Case Study of Simple Function Classesarxiv.org
  3. An Explanation of In-context Learning as Implicit Bayesian Inferencearxiv.org
  4. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?arxiv.org
  5. Transformers Learn In-Context by Gradient Descentproceedings.mlr.press
  6. Calibrate Before Use: Improving Few-Shot Performance of Language Modelsarxiv.org
  7. Lost in the Middle: How Language Models Use Long Contextsarxiv.org