Techniques

Chain-of-thought prompting and self-consistency

Chain-of-thought prompting asks a model to reason step by step before answering. When it helps, when reasoning models make it unnecessary, and self-consistency.

Chain-of-thought (CoT) prompting asks a model to write out its reasoning before it gives a final answer. The idea was described by Wei and colleagues at Google in 2022, who showed that worked examples with intermediate steps improved results on arithmetic and logic problems.

Why writing the steps helps

A language model produces one token at a time, each conditioned on everything before it. If it must answer immediately, the answer is predicted with no intermediate work. If it first writes the steps, those steps become part of the context the answer is predicted from. Reasoning on the page is working memory.

Zero-shot chain-of-thought

The simplest form adds an instruction to reason first.

A warehouse ships 240 orders a day. 15% need gift wrap, which adds
4 minutes each. How many staff hours a day does gift wrap take?

Work through the problem step by step, then give the final answer
on its own line as "Answer: <number> hours".

Two details matter. Ask for the reasoning before the answer, and give the answer a fixed marker so your code can find it.

Few-shot chain-of-thought

Here the examples show the reasoning as well as the result, so the model imitates the method.

Q: A plan costs $12 a month or $120 a year. How much does the yearly
plan save over 2 years?
Reasoning: Monthly for 2 years is 12 x 24 = $288. Yearly for 2 years
is 120 x 2 = $240. The difference is 288 - 240 = $48.
Answer: $48

Q: {{question}}
Reasoning:

Use this when the task has a particular procedure you want followed.

Structured reasoning

For answers that code will read, separate thinking from the result with tags or fields.

Think through the eligibility rules inside <reasoning> tags.
Then give the decision inside <decision> tags as "approve" or "reject".

In JSON, put the reasoning field before the verdict field. Fields are generated in order, so reasoning that comes after the verdict cannot influence it. See structured output prompting.

Reasoning models change the advice

Many current models reason internally before they respond, often under a setting called thinking, reasoning effort, or extended thinking. For these models:

  • Do not add “think step by step”. They already do, and the instruction can make output longer without making it better.
  • Describe the goal and the constraints, and let the model choose its approach.
  • Use the provider’s control for how much reasoning to spend, instead of prompt wording.
  • Still ask for a brief justification in the output when a person needs to check the answer.

Explicit chain-of-thought remains useful with smaller and faster models that do not reason internally, and whenever you need the reasoning visible in the response.

When chain-of-thought helps and when it does not

Task Helps?
Multi-step arithmetic or logic Yes
Decisions with several rules to apply Yes
Debugging: finding the cause from symptoms Yes
Simple classification or extraction Rarely; it adds cost
Creative writing No
Lookup of a fact in supplied text No

The cost is real: reasoning is output tokens, which are usually priced higher than input tokens and add latency.

A caution about explanations

The written reasoning is a plausible account, not a guaranteed trace of how the answer was produced. A model can reach a conclusion and write a justification that sounds right. Treat the reasoning as something to verify, especially where a decision affects people.

Self-consistency

Self-consistency, described by Wang and colleagues in 2022, builds on chain-of-thought. Run the same reasoning prompt several times with sampling enabled, so the paths differ, then take the most common final answer.

  1. Send the CoT prompt N times, for example 5, with a temperature above zero.
  2. Extract the final answer from each response.
  3. Return the answer that appears most often.

Different reasoning paths make different mistakes, but correct paths tend to agree. The vote filters out the occasional slip.

Use it when the problem has one correct answer that can be compared exactly (a number, a label, a yes or no), and the accuracy is worth N times the cost.

Skip it when answers are free text that cannot be voted on, or when a single call is already reliable. If the votes are split evenly, treat that as a signal that the question is ambiguous or the prompt is underspecified.

From reasoning to pipelines

When the steps of a problem are known in advance, do not ask one prompt to reason through all of them. Give each step its own prompt. That is prompt chaining. When the steps depend on information the model must fetch, combine reasoning with tools: that is ReAct prompting.

Check the prompt first

Reasoning cannot fix an unclear question. A CoT prompt with a vague objective or conflicting rules reasons carefully toward the wrong answer. Run it through the PromptFlowEngine prompt analyzer before you spend tokens on reasoning.