Prompt engineering is the practice of designing, testing, and refining inputs to generative artificial intelligence systems, especially large language models, to obtain outputs that satisfy specified requirements. A prompt can contain instructions, questions, examples, reference material, and conversation history. Rather than simply choosing persuasive wording, the practice involves defining tasks, organizing information, specifying output constraints, and evaluating results. Its effectiveness depends on the model, the task, and the surrounding application. (developers.openai.com)
Technical foundations
A language model generates text conditioned on its input. Prompt engineering changes that conditioning without ordinarily changing the model’s learned parameters. This distinguishes it from fine-tuning, which updates parameters using task-specific data. Prompts can therefore adapt a pretrained model to different tasks without a separate training procedure for every application, although adaptation through prompting does not guarantee reliable performance. (arxiv.org)
In text-based systems, tokenization converts input into units processed by the model. The context window limits how much material can be available in a generation request. Instructions, examples, reference documents, and conversation history compete for this capacity. A larger window does not necessarily mean that every included detail will be used equally well: research has documented sensitivity to where relevant information appears in long inputs. (developers.openai.com)
Chat interfaces may distinguish messages by role, separating application-level instructions from user requests and other content. Prompt design can exploit these distinctions, but role semantics and instruction priorities depend on the implementation. Structured boundaries, such as headings or tagged sections, help identify which text is an instruction, an example, or source material. (developers.openai.com)
Development and research background
An important milestone was the 2020 paper Language Models are Few-Shot Learners. It demonstrated that GPT-3 could perform numerous natural language processing tasks from instructions and examples supplied in its input, without task-specific gradient updates. The experiments included machine translation, question answering, and other language tasks. This helped establish in-context learning as an important method of using pretrained models. (arxiv.org)
Research subsequently investigated more elaborate prompting strategies. A 2022 study introduced chain-of-thought prompting, in which demonstrations include intermediate reasoning steps rather than only final answers. It reported improvements on arithmetic, commonsense, and symbolic reasoning benchmarks for sufficiently large models. These findings describe particular experimental settings, not a universal rule that longer explanations always improve results. (arxiv.org)
A related but distinct approach is prompt tuning. Instead of manually writing ordinary text, this method learns continuous prompt representations while keeping the underlying model frozen. It belongs to parameter-efficient adaptation and involves optimization during training, unlike ordinary text-prompt revision. (arxiv.org)
Prompt structure and techniques
A task-oriented prompt commonly specifies the intended operation, relevant context, constraints, and expected response format. Instructions may define an audience, desired level of detail, classification labels, or rules for handling missing information. Explicit requirements reduce ambiguity, while delimiters separate instructions from material to be analyzed. Assigning a role can establish a perspective or style, but does not confer expertise or independently verify the output. (docs.anthropic.com)
Zero-shot prompting supplies a task without demonstrations. Few-shot prompting includes example inputs and desired outputs, allowing the model to infer patterns such as label meanings, formatting, or transformation rules. Demonstrations guide behavior within the current context rather than becoming permanent additions to the model’s training data. (arxiv.org)
For multi-step tasks, prompting may separate evidence extraction, analysis, and final presentation. Prompt chaining distributes these operations across multiple calls, passing one stage’s output into another. Such workflows make intermediate results available for inspection, but their design must account for mistakes carried from one stage to the next. The usefulness of explicit reasoning instructions varies by model; prompting guidance for reasoning-oriented systems can differ from guidance for other models. (docs.anthropic.com)
External information and application integration
Retrieval-augmented generation connects generation with information retrieved from an external collection. Retrieved passages become evidence available to the model, supporting tasks that require knowledge beyond the immediate user request. Prompt engineering governs how that evidence is presented and how the model is instructed to use it; retrieval itself is a separate component. (arxiv.org)
In applications using a knowledge base or tools, prompt design also defines workflow conditions: which information is needed, when an external operation is appropriate, and how results enter the response. This makes prompting one part of system design rather than a substitute for retrieval quality, tool implementation, or output checking. (docs.anthropic.com)
Evaluation and reproducibility
Prompt quality is assessed against task-specific criteria, such as answer correctness, format compliance, completeness, or successful execution. Evaluation combines representative examples with difficult and unexpected cases. Human judgments and automated checks serve different purposes, while model-generated grading requires scrutiny of its own reliability. (developers.openai.com)
Repeated revision on the same examples can create overfitting to an evaluation set. Separating development examples from a held-out test set provides a stronger check of generalization. Because outputs can vary and model updates can change behavior, reproducible comparisons record prompt versions, model versions, and relevant generation settings. Evaluation is repeated when these components change. (developers.openai.com)
Limitations and security
Prompting does not eliminate AI hallucinations, and supplying reference material does not ensure that the model will use it correctly. Evidence placement can affect performance, as shown in long-context experiments. Fluent responses therefore remain distinct from verified answers. (arxiv.org)
Prompt injection is a security problem in which adversarial instructions embedded in processed content attempt to redirect a system’s behavior. External webpages, documents, and tool results can carry such instructions. Prompt boundaries are useful, but defenses also involve model training, content detection, application safeguards, and testing; prompt wording alone does not make an agent immune to attack. (anthropic.com)