Plan-and-Solve

This page follows Wang et al. (2023), Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models (arXiv:2305.04091v3, ACL 2023). PS is best read as a targeted repair to zero-shot CoT, not as a general planning system.

Paper Question

Zero-shot CoT uses a simple trigger such as “Let’s think step by step.” The paper asks why it still fails and whether a better zero-shot trigger can reduce those failures without demonstrations or finetuning.

The authors manually analyze 100 arithmetic examples and identify three error types:

  • calculation errors;
  • missing-step errors;
  • semantic misunderstanding errors.

Method

Plan-and-Solve changes the trigger so the model first understands the problem, devises a plan, then solves according to that plan.

PS targets missing-step errors. PS+ adds more explicit instructions to extract relevant variables, calculate intermediate results, and pay attention to calculation and commonsense. PS+ targets calculation errors as well.

The method remains zero-shot and single-sample. There are no examples, tools, verifiers, retries, or search branches.

Evidence

The paper evaluates ten datasets:

  • six arithmetic datasets: GSM8K, SVAMP, MultiArith, AddSub, AQuA, SingleEq;
  • two commonsense datasets: CommonsenseQA and StrategyQA;
  • two symbolic datasets: Last Letter and Coin Flip.

On the six arithmetic datasets, PS+ averages 76.7, improving over zero-shot CoT and approaching few-shot manual CoT. It also improves CommonsenseQA, StrategyQA, Last Letter, and Coin Flip in the reported experiments.

The error analysis is the strongest evidence for the paper’s mechanism: zero-shot CoT has 7% calculation, 12% missing-step, and 27% semantic errors in the sampled analysis; PS+ reduces calculation and missing-step errors while semantic misunderstanding remains unchanged.

What The Prompt Ablations Show

The planning instruction is doing real work. Adding detailed execution instructions without an explicit plan does not explain the gain. The model needs to externalize the task structure before executing the steps.

That makes PS a minimal planner-executor pattern: plan first, solve second. It is still a prompt, but it foreshadows harness designs where planning becomes a separate role, state, or checkpoint.

Limits

The paper’s own limitation is clear: semantic misunderstanding remains. If the model reads the problem wrong, a plan can make the wrong interpretation more coherent.

PS is also a single-chain method. It has no observation, no verifier, no retry, and no backtracking. A wrong plan can propagate through the whole answer.

The experiments use GPT-3-era models, so the marginal benefit may change as models become stronger or are trained to plan by default.

Agent Design Value

My design takeaway is that “plan before acting” is a cheap guard against omitted steps, but a weak guard against wrong understanding.

For agents, PS suggests a useful checkpoint: before executing a complex task, force the model to state the plan in a form that can be inspected, edited, or rejected. But do not confuse a plan with validation. The plan should become input to verification, tool readiness checks, permission checks, and progress tracking.

In harness terms, PS is the smallest seed of lifecycle orchestration. It becomes serious only when the plan is connected to execution state, observations, failure recovery, and evaluators.

Was this page helpful?