Prompt chaining means solving a task with several prompts in sequence, where the output of one becomes the input of the next. Prompt decomposition is the design step before it: deciding how to split the task. Together they are the most dependable way to make a model do something complicated.
Why one big prompt fails
A prompt that asks a model to extract, analyze, decide, and write in one pass has several problems:
- Instructions compete. The more a prompt asks for, the more likely one requirement is neglected.
- You cannot tell which part failed. A wrong final answer could come from any step.
- You cannot test the parts. There is nothing to assert on until the end.
- Every call pays for every instruction, even those irrelevant to this input.
A chain gives each step one job, a clear input, and a checkable output.
How to decompose a task
Ask what a careful person would do, in order, and where they would hand work to someone else. Good split points are where:
- The kind of work changes (reading, then judging, then writing).
- An intermediate result is worth checking (the extracted facts, before conclusions are drawn).
- A step only sometimes applies (translate only if the text is not English).
- Different steps need different models (a small model to classify, a larger one to write).
A worked example
The task: answer a customer complaint email.
Step 1: extract. A small, fast model turns the email into structured facts.
Extract from the email in the tags: order_id, product, the problem in
one sentence, and what the customer is asking for. Return JSON only.
Use null for anything not stated.
<email>
{{email}}
</email>
Step 2: decide. Code looks up the order. A prompt applies the policy.
Given the complaint and the order record, choose one action from:
refund, replace, escalate, explain. Give the policy clause that applies.
Return JSON: {"action": string, "clause": string}.
<complaint>{{step1_json}}</complaint>
<order>{{order_record}}</order>
<policy>{{policy}}</policy>
Step 3: write. A third prompt drafts the reply from the decision.
Write a reply to the customer, under 120 words, that states the action
and what happens next. Do not promise anything outside the action.
<action>{{step2_json}}</action>
<customer_name>{{name}}</customer_name>
Each step can be tested alone. If replies promise the wrong thing, you look at step 2’s output and know immediately whether the decision or the writing is at fault.
Passing data between steps
Use structured output between steps. JSON or tagged fields, not prose. The next prompt, and your code, can then rely on the shape. See structured output prompting.
Pass only what the next step needs. Forwarding the whole history makes later prompts long and lets early mistakes spread.
Validate at every boundary. Check the output of each step before using it. A chain without validation multiplies errors: five steps that are each right 95% of the time are all right only about 77% of the time.
Delimit everything you insert. The output of one model is untrusted input to the next. A complaint email that says “ignore your instructions” has passed through step 1 and is now inside step 2’s prompt. Wrap inserted content in tags. See prompt injection.
Chain shapes
| Shape | Description | Example |
|---|---|---|
| Sequential | A then B then C | Extract, decide, write |
| Routing | Classify, then pick a branch | Billing questions go to one prompt, bugs to another |
| Parallel | Independent steps at once, then merge | Summarize 10 documents, then combine |
| Generate and check | One prompt produces, another reviews | Draft, then critique against the requirements |
| Loop | Repeat until a condition holds | Revise until validation passes, with a retry limit |
Always cap loops. A chain that retries forever is an outage and an invoice.
Chaining versus the alternatives
Versus chain-of-thought. Chain-of-thought keeps the reasoning inside one response. Use it when the steps are not known ahead of time. Use a chain when they are: you get testability and control.
Versus agents. In a chain, your code decides the order. In an agent, the model decides. A chain is cheaper, faster, and predictable. Choose an agent only when the path depends on what is discovered along the way. See ReAct prompting and prompts for AI agents.
Costs to watch
A chain makes more calls. Keep it economical:
- Use the smallest model that handles each step.
- Run independent steps in parallel to cut latency.
- Skip steps that do not apply to this input.
- Cache the fixed part of each prompt where the provider supports it. See token optimization.
Keep each link healthy
A chain has several prompts to maintain, so a problem in one is easy to miss. Store each as a file and check them all on every change. The prompt testing gate fails a pull request when any prompt in the repository drops below your quality bar or exceeds its token budget, and the prompt analyzer reports undelimited variables, which matter most where one step feeds the next.