Chain-of-Thought
This page follows Wei et al. (2022), Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (arXiv:2201.11903v6, NeurIPS 2022). The point is not “ask the model to explain itself.” The paper’s claim is narrower and more important: for sufficiently large models, placing natural-language intermediate reasoning before the answer elicits capabilities that standard input-output prompting leaves hidden.
Paper Question
The paper asks whether large language models can solve multi-step reasoning tasks using only prompting. Prior ways to improve arithmetic or symbolic reasoning often required finetuning, external verifiers, formal programs, or specialized architectures. CoT tests a cheaper intervention: change the few-shot exemplars so each answer is preceded by a reasoning chain.
The target tasks are deliberately broad:
- arithmetic word problems such as GSM8K, SVAMP, ASDiv, AQuA, MAWPS;
- commonsense reasoning such as CSQA, StrategyQA, date, sports, and SayCan-style tasks;
- symbolic manipulation such as last-letter concatenation and coin-flip state tracking.
Method
Standard few-shot prompting uses examples shaped like:
Question -> Answer
Chain-of-Thought prompting uses examples shaped like:
Question -> natural-language intermediate reasoning -> Answer
The method does not update weights, call tools, or add a verifier. In the arithmetic experiments, the paper uses the same eight human-written CoT exemplars across datasets. The model is expected to infer the format and generate its own chain before its final answer.
This matters because the final answer is conditioned on the generated intermediate steps. CoT turns a single hard jump into a sequence of smaller generation decisions.
Evidence
The headline result is GSM8K. PaLM 540B moves from 17.9% with standard prompting to roughly 57% with CoT prompting, surpassing a then state-of-the-art finetuned system with a verifier. The paper also shows gains across other arithmetic datasets, commonsense tasks, and symbolic tasks.
Three findings define the paper:
- CoT is emergent with scale. The gains are absent or negative in smaller models and become strong around the largest model scale studied.
- The benefit concentrates on multi-step tasks. Easy single-operation arithmetic does not benefit much; hard multi-step GSM8K examples benefit most.
- The effect is robust to many exemplar changes. Different annotators, shorter chains, different exemplar subsets, and exemplar order changes still preserve the core advantage.
What The Ablations Show
The paper’s ablations prevent an overly shallow interpretation:
- Equation-only prompting does not replace natural-language reasoning on semantically hard problems.
- Extra tokens alone do not help; placeholder symbols with similar length perform like the baseline.
- Putting the chain after the answer does not help, which suggests the chain must influence the answer rather than merely decorate it.
So the active ingredient is not verbosity, hidden compute, or explanation style. It is answer-before-chain versus chain-before-answer.
Limits
CoT does not prove that the written reasoning is faithful to the model’s internal computation. The paper manually inspects examples and finds many correct answers have plausible chains, but it does not claim causal faithfulness.
CoT also has no recovery mechanism. It is one left-to-right sample. If the model commits to a wrong intermediate premise, there is no observation, verifier, search branch, or retry loop to correct it.
The paper also makes scale a first-class limitation: small models can produce fluent invalid chains. CoT amplifies latent reasoning ability; it does not create that ability from nothing.
Agent Design Value
My design takeaway is that CoT should be treated as a trace primitive, not as a truth primitive.
For agents, explicit reasoning is useful because it exposes intermediate intent, assumptions, subgoals, and possible failure points. That helps observability and human intervention. But the trace is not evidence of success. A production harness should let the model reason, then verify outcomes through tools, tests, environment state, or human review.
CoT becomes the inner language of later methods: ReAct inserts observations between reasoning steps; Reflexion writes lessons after failed traces; Plan-and-Solve structures the first part of the chain as a plan; LATS searches over many possible chains and actions.