Reasoning & Planning

In the process of building autonomous AI agents, if the large language model (LLM) is the agent's brain and tool use is its hands and feet, thenReasoning & Planningis the core engine that upgrades it from a simple Q&A machine to an autonomous problem solver.

Complex real-world tasks often cannot be completed through one-pass generation. AI needs the ability to decompose goals, perform logical reasoning, explore paths, self-correct, and orchestrate tools.

The following are the most mainstream reasoning and planning frameworks in the industry today.


Chain of Thought (CoT)

Step-by-step reasoning capability

Traditional LLMs often generate answers intuitively in one step.

The core idea of Chain of Thought (CoT) is:Forcing the model to explicitly output intermediate reasoning steps (Let's think step by step). This approach can significantly activate the model's potential in complex math, logical reasoning, and common-sense Q&A.

CoT not only gives the model more computation time (token count represents computational effort), but also allows subsequent generation to build on the correct logic from earlier steps.

Example: Few-shot CoT prompt design

# Guide the model to perform CoT reasoning by providing examples that include the reasoning process.
prompt = """
Problem: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have in total?
Solution: Roger initially has 5 tennis balls. 2 cans of tennis balls, 3 each, give 2 * 3 = 6 tennis balls. 5 + 6 = 11. The answer is 11.

Problem: The cafeteria has 23 apples. If they use 20 for lunch and buy 6 more, how many apples are there now?
Solution: The cafeteria originally had 23 apples. After using 20, 23 - 20 = 3 remain. After buying 6 more, there are now 3 + 6 = 9. The answer is 9.

Problem: {user_question}
Solution: """


ReAct framework (Reasoning + Acting)

Reasoning + Acting loop

If CoT is just building a cart behind closed doors inside the model, thenReAct(Reason + Act)it is letting the model open its eyes and see the world. It interweaves internal logical reasoning (Thought) with external tool interaction (Action) to form a dynamic closed-loop feedback system.

Under the ReAct paradigm, the Agent followsThought -> Action -> Observationloop until reaching a final conclusion.

ReAct Loop Architecture Diagram Shows the loop relationship among Thought, Action, and Observation. Thought Analyze the current state and goals Action Call external tools/APIs Observation User Query

Limitations:ReAct performs excellently in short-term tasks with clear steps. However, since the entire history of thoughts and actions is packed into the same context window, when the task chain is too long, it is very easy to fall into an infinite loop or forget the initial goal due to context overload.


Plan-and-Execute (plan-first execution mode)

To address ReAct's weakness in long-horizon tasks, Plan-and-Execute decouples thinking and acting, adopting a strategy similar to how humans handle large projects:First create a schedule, then execute tasks one by one.。

The system is usually divided into two independent roles:

  1. Planner: Responsible for receiving the overall goal and generating a detailed step-by-step subtask list.
  2. Executor: Responsible for executing these subtasks in order. The executor is usually a small ReAct Agent that focuses on completing only the current small goal each time.
Plan-and-Execute Architecture Diagram Complex Goal Planner Scrape web pages Extract data Generate report Executor Process subtasks one by one (Each with independent context) Final delivery

Tree of Thoughts (ToT) and tree-like multi-path exploration

Tree-like multi-path exploration

Whether it is CoT or Plan-and-Execute, it is essentiallylinearpath exploration. However, when writing code, solving math problems, or doing creative writing, humans often envision multiple options, evaluate them, choose the best one, and even backtrack when errors are found.

ToT (Tree of Thoughts)It models the reasoning process as a tree: nodes are the current thought states. At each branching point, the model generates multiple candidate Thoughts, then uses an internal Evaluator to score these nodes (e.g., feasible, potentially risky, infeasible). Combined withBFS (Breadth-First Search)orDFS (Depth-First Search)algorithm, it decides whether to continue deeper or backtrack and retry.


Task planning & MCTS (Monte Carlo Tree Search)

Complex task decomposition and search

When dealing with strategy games or extremely difficult reasoning tasks (such as frontier mathematical verification and complex codebase refactoring), simple ToT is still not efficient enough. The industry has begun combining LLMs with traditional reinforcement learning search algorithmsMCTS(Monte Carlo Tree Search)Combining them (similar to AlphaGo's core logic).

  • LLM as a Policy Network: Provides heuristic next-step action suggestions, reducing meaningless branch expansion.
  • LLM/code environment as a Value Network: Through simulated execution (Rollout), it predicts the final win rate or success rate of an action sequence.
  • Advantages: In a vast solution space, it can find the planning path with the greatest potential for global optimality.

Reflexion: self-reflection and error correction

When humans execute tasks, if they fail on the first attempt, they summarize lessons learned and avoid mistakes in the next attempt.

ReflexionThis framework gives the Agent similar capabilities.

In the Reflexion closed loop, when the Agent's output is judged as a failure (e.g., test cases fail, API errors occur), it triggers aReviewer mechanism。

The LLM is asked to write a colloquial reflection (Reflection) based on historical actions and failure feedback, for example:"I just used the wrong API parameter format. Next time, I should consult the documentation before passing JSON."This reflection will be stored in episodic memory as contextual prompts for the next attempt, thereby greatly enhancing the agent's self-healing capabilities.

Example: Reflexion's reflection prompt design

reflection_prompt = """
You are an AI assistant trying to write a Python crawler.
This is the code you just executed: {previous_code}
This is the error message returned by the runtime environment: {error_traceback}

Please reflect deeply:
1. What is the root cause of the error?
2. What is your specific modification strategy for the next attempt?

Please record the reflection to guide subsequent actions.
"""


Task decomposition strategies and engineering practices

In real production-grade AI agent development, relying purely on LLM "zero-shot" for complex planning is unstable. Common hybrid intervention strategies include:

Intervention Strategy Core Approach Applicable Scenarios
Subtask Templating (SOP) Instead of letting the LLM plan freely, predefined standard operating procedures (SOPs) are used to make the LLM flow within a fixed state machine. Customer service systems, standardized data cleaning pipelines.
HITL (Human-in-the-Loop) After the Planner generates the task list, interrupt execution and require a human user to confirm, modify, or approve, then hand it over to the Executor for execution. High-risk operations: such as deleting database records, sending mass emails, and large fund transfers.
RLHF-guided Planning Use reinforcement learning and human preference feedback to fine-tune the planning capability of large models, making them more inclined to generate safe and efficient step combinations. The training stage of the underlying large language model base (e.g., OpenAI's o1 model training).

Framework comparison summary

Pattern Core Mechanism Advantages Disadvantages
CoT Step-by-Step Linear Reasoning Minimal implementation, significantly improves basic reasoning accuracy Cannot call external tools, easily goes down a dead-end path
ReAct Alternating Thinking and Action Loop Dynamically adapts to the environment, can adjust in real time through observations Context easily explodes as steps accumulate, losing sight of the original goal
Plan-and-Execute Decompose into subtasks first, then execute in isolation Extremely suitable for long-horizon complex tasks, clear context Not flexible enough when facing sudden changes (when the plan itself is wrong)
ToT / MCTS Tree search, evaluation and backtracking Can solve the most difficult complex logic problems The computational cost is extremely high, and token consumption is exponential.
Reflexion Generate reflective memory based on failure feedback. Possess the ability to self-correct and continuously evolve. Rely on explicit feedback signals (such as code compiler errors)
Other extensions