Reflexion

This page follows Shinn et al. (2023), Reflexion: Language Agents with Verbal Reinforcement Learning (arXiv:2303.11366v4, NeurIPS 2023). Reflexion moves learning from parameter space into context space.

Paper Question

The paper asks how a language agent can improve from trial and error when gradient updates are unavailable or too expensive. Traditional RL converts scalar reward into weight updates. Reflexion converts feedback into language and stores it in memory for the next attempt.

This is why the paper calls the method verbal reinforcement.

Method

Reflexion has three model roles:

  • Actor: an LLM policy that generates text and actions, often using CoT or ReAct.
  • Evaluator: a task-specific scorer for the trajectory, such as exact match, a heuristic, an LLM judge, generated tests, compiler output, or environment success.
  • Self-Reflection model: an LLM that reads the trajectory and feedback, then writes a natural-language lesson.

Memory has two parts:

  • short-term memory: the current trajectory;
  • long-term memory: reflections from prior attempts.

The loop is:

Actor attempts task -> Evaluator scores trajectory -> Self-Reflection writes lesson -> lesson enters memory -> Actor retries

The policy is effectively {LLM weights, memory}. The weights stay fixed; memory changes.

Evidence

The paper evaluates three task families:

  • sequential decision-making in ALFWorld;
  • reasoning in HotpotQA;
  • programming in HumanEval, MBPP, Rust translations, and LeetcodeHardGym.

The reported headline gains are 22% on ALFWorld, 20% on HotpotQA, and 11% on HumanEval. The paper reports 91% pass@1 on HumanEval in its Reflexion setup.

The coding setup is especially revealing. Generated unit tests give the evaluator a concrete feedback signal. Reflection then turns that signal into a repair instruction. Where tests are noisy, as in MBPP Python in the paper’s analysis, Reflexion can be misled by false positives.

What The Ablations Show

The key result is not that “retrying helps.” The paper shows that retrying without reflection is weak.

In reasoning tasks, ReAct-only, CoT-only, and CoT-with-ground-truth-context baselines rarely recover failed examples across later trials. Episodic memory helps, but self-reflection over the trajectory adds more.

In programming, removing test generation removes the agent’s ability to know whether the attempt worked; removing self-reflection removes the mechanism that turns failure into a next-attempt strategy.

So the causal structure is: evaluator signal -> diagnosis -> memory -> changed next attempt.

Limits

Reflexion is only as good as the evaluator. If feedback is wrong, the agent may write a confident false lesson and become worse. The method also assumes multiple attempts are allowed and that attempts can be compared against a stable task.

Memory capacity is small in the original setup. Long-running agents need retrieval, summarization, or structured memory to avoid filling context with stale or low-value lessons.

The method has no formal success guarantee and can get trapped in local minima, especially in tasks where the required next behavior is very different from previous failed behavior.

Agent Design Value

My design takeaway is: retries are cheap only when they are diagnostic.

A production harness should not merely “try again.” It should capture the failed trace, score it with the strongest available evaluator, write a compact lesson, and carry only useful lessons forward. That makes Reflexion a bridge between Context, Observability, and Verification: the memory should store failure-reducing knowledge, not a transcript.

Reflexion also makes model improvement inspectable. Unlike gradient updates, the policy update is a readable sentence. That is valuable for debugging and governance, but also risky: if the lesson is wrong, it becomes a persistent steering instruction.

Was this page helpful?