ReAct

This page follows Yao et al. (2022), ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629v3, ICLR 2023). ReAct is the paper that turns CoT from a static reasoning trace into an interactive agent loop.

Paper Question

The paper starts from two failures:

  • CoT-style reasoning can hallucinate because it reasons without fresh external evidence.
  • Act-only systems can use tools but lack language-level planning, working memory, and abstraction.

ReAct asks whether a language model can interleave both: reason in natural language, act in an environment, observe the result, and continue.

Method

The formal move is simple. The action space changes from A to A union L, where A is external actions and L is language. A language action is called a thought. It does not change the external environment, but it updates the trajectory context.

The loop is:

Thought -> Action -> Observation -> Thought -> Action -> Observation -> ... -> Finish

In HotpotQA and Fever, actions are Wikipedia API calls such as search, lookup, and finish. In ALFWorld and WebShop, actions are environment commands. Thought density is task-dependent: knowledge tasks use dense thought-action-observation steps, while decision tasks can use sparser thoughts around important subgoals.

Evidence

The paper evaluates four environments:

  • HotpotQA and Fever for knowledge reasoning with Wikipedia access;
  • ALFWorld for text-game decision-making;
  • WebShop for web shopping navigation.

The knowledge results are deliberately nuanced. ReAct beats act-only prompting and beats CoT on Fever, but slightly lags CoT on HotpotQA exact match. The best systems combine ReAct and CoT self-consistency: if ReAct fails within a step budget, fall back to CoT-SC; if CoT-SC lacks majority confidence, fall back to ReAct.

The decision-task results are stronger. On ALFWorld, the best ReAct run reaches 71% success, beating the best Act run and the BUTLER imitation-learning baseline. On WebShop, ReAct improves over imitation and RL baselines in the paper.

What The Error Analysis Shows

The most important table is the HotpotQA failure analysis:

  • CoT failures are dominated by hallucination.
  • ReAct almost eliminates that hallucination category because observations ground the trajectory.
  • ReAct introduces new failures: bad retrieval, repetitive loops, and reasoning errors caused by the rigid interleaving.

This is the paper’s real lesson: grounding is not free. External observations reduce hallucination, but tool failure becomes part of the reasoning process. ReAct is not “tools beat thinking”; it is “thinking and tools repair different failure modes.”

The finetuning result is also important. Prompt-only ReAct is hard for smaller models, but with 3,000 correct trajectories finetuning makes ReAct the best of the compared methods. That suggests ReAct is not only a prompting pattern; it is also a trajectory data format.

Limits

ReAct is a single-trajectory method. It can revise within a run, but it does not learn across attempts unless another memory mechanism is added. It also does not preserve multiple branches or backtrack.

It is bounded by the action space and observation quality. If search fails, the environment lies, or the tool surface is too brittle, the reasoning loop can become confidently grounded in the wrong evidence.

The paper also avoids dangerous real-world actions. WebShop is a benchmark environment; the agent is not actually buying products. Production systems need governance around the actions that ReAct makes easier to invoke.

Agent Design Value

My design takeaway is that ReAct is the minimum viable lifecycle loop for tool agents.

The non-negotiable implementation rule is: every action must produce an observation that is fed into the next reasoning step. If a harness hides observations, compresses them too aggressively, or lets the model continue as if an action succeeded, it loses the main benefit of ReAct.

ReAct also explains why traces should separate thought, action, and observation. That separation lets humans and evaluators ask: did the model reason badly, choose the wrong tool, receive bad evidence, or ignore a good observation? That is exactly the bridge from model-side reasoning to ETCLOVG’s lifecycle, observability, verification, and governance layers.

Was this page helpful?