Prompt engineering is the practice of designing, testing, and refining the instructions given to a large language model (LLM) so that it produces the output you need, reliably. The word “engineering” matters: the goal is not one good answer, it is the same quality of answer every time the prompt runs.
A working definition
A prompt is everything the model receives before it starts writing: instructions, background, examples, documents, and the question itself. Prompt engineering is deciding what goes into that input and how it is arranged.
It covers three activities:
- Specifying the task so nothing important is left to guesswork.
- Supplying context the model does not have, such as your data, your audience, and your constraints.
- Measuring the result and changing the prompt based on evidence, not on a hunch.
How prompt engineering works
A language model predicts the next token from the tokens it has seen so far. It has no access to your intent, only to your text. Two consequences follow.
Anything you leave out gets filled in. Ask for “a short summary” and the model picks a length, an audience, and a format. It will pick differently next time. Most inconsistent output is caused by a decision the prompt never made.
Position and structure carry meaning. A model treats a heading, a delimiter, or an example as a signal about what kind of text comes next. Separating instructions from data, and showing the shape of the answer you want, changes the prediction more than adding adjectives does.
Here is the same request written twice.
Summarize this report.
Summarize the quarterly sales report below for the EMEA leadership team.
Write 5 bullet points covering revenue by country and the change from
last quarter. Return a markdown list.
<report>
{{report}}
</report>
The second version names the audience, the length, the content, and the format, and it marks where the data starts and ends. Nothing in it is clever. It simply makes the decisions that the first version left open.
Why prompt engineering matters
Reliability. Software that calls a model needs predictable output. A prompt that returns JSON nine times out of ten breaks the tenth request.
Cost. Models are billed by the token. A prompt that repeats itself, or carries context it does not need, costs more on every call. See token optimization for how to measure and reduce it.
Safety. Prompts often include text from users, documents, and web pages. Without care, that text can override your instructions. This is called prompt injection, and defending against it starts with how the prompt is built.
Speed of iteration. Changing a prompt takes minutes. Fine-tuning a model takes data, money, and time. For most tasks, a better prompt is the cheapest improvement available.
What prompt engineering is not
It is not a list of magic phrases. Tricks such as promising a tip or writing in capitals have inconsistent effects that change between model versions. What lasts is clear specification, relevant context, and testing.
It is also not a replacement for knowing the task. A prompt can only ask for what its author understands well enough to describe.
Prompt engineering and context engineering
As models became able to read very large inputs and call tools, the work shifted from wording a single instruction to deciding what information the model sees at each step. This is often called context engineering. It includes retrieval, agent instructions, and tool connections such as the Model Context Protocol. The principles are the same: give the model what it needs, clearly separated, and nothing that distracts it.
Where to go next
- Follow the step-by-step prompt engineering guide to write your first production prompt.
- Learn the types of prompts and the main prompt engineering techniques.
- Paste a prompt into the PromptFlowEngine analyzer to see which decisions it leaves open.