LATS

This page follows the existing foundations arc around LATS: Language Agent Tree Search Unifies Reasoning, Acting and Planning in Language Models (Zhou et al., ICML 2024). The local tmp/arxiv folder does not currently include this PDF, so this page is kept to the algorithmic facts already represented in the Chinese foundations draft.

Method

Earlier paradigms usually generate one trajectory from left to right. ReAct walks one thought-action-observation path. Reflexion can retry after failure, but each attempt is still mostly linear. LATS adds deliberate search.

LATS places partial agent trajectories in a Monte Carlo Tree Search framework. Each node is a partial trajectory from the root to the current state. Each iteration performs:

  1. Selection: choose promising nodes with an upper-confidence rule.
  2. Expansion: sample several ReAct-style next steps.
  3. Evaluation: score new states with an LM value estimate and self-consistency signal.
  4. Simulation: roll out a promising path to a terminal state.
  5. Backpropagation: update values and visit counts along the path.
  6. Reflection: when a trajectory fails, write a Reflexion-style lesson for later search iterations.

What It Combines

LATS makes the previous methods composable:

MethodRole inside LATS
CoTReasoning inside each node.
ReActThe step primitive: reason, act, observe.
ReflexionLessons from failed rollouts.
MCTSThe outer structure that supports branching, value-guided selection, and backtracking.

The important shift is from one left-to-right trajectory to an explicit search tree.

Engineering Meaning

Search is useful only when the environment can support it. Pure reasoning tasks, code sandboxes, and resettable simulators are good fits because failed branches can be abandoned or replayed. Irreversible actions are not: once an email is sent, money is transferred, or a database is mutated, “backtracking” is no longer a harmless search operation.

The value function is the load-bearing part. Without a useful value signal, search becomes expensive enumeration. This connects LATS directly to the verification layer: tests, compilers, formal checks, and environment success signals make search far more valuable than weak self-evaluation alone.

Value For Agent Design

Use LATS-like search for high-value tasks where extra latency and token cost are justified, where candidate paths can be evaluated, and where state can be reset or safely simulated. Keep side effects outside the search tree until the system has committed to an action path.

LATS is also a design pattern for harness engineering: reasoning, acting, reflection, state reset, evaluation, and selection are independent mechanisms that can be composed into a stronger controller.

Was this page helpful?