AI Agent Core Components

If we compare an AI Agent to asmart restaurant, how does it turn your needs into dishes served to you? This relies on its four core components:Brain, Tools, Memory, and Planning。

  • Brain:It is responsible for understanding orders, determining goals, and deciding the sequence; it is the command center of the restaurant.
  • Tools:It is responsible for hands-on execution, including cutting and preparing, cooking, purchasing, and other actions, turning decisions into executable operations.
  • Memory:It is responsible for recording customer preferences, current steps, and processed content, ensuring the process is not chaotic or repetitive.
  • Planning:It is responsible for breaking the whole dish into steps, determining the order, and ensuring tasks progress through the process to completion.


Overall Architecture

The following figure shows the five layers of AI Agent components and their collaboration. The perception layer receives external input, the brain is responsible for understanding and decision-making, the planning layer decomposes tasks, the tool layer is responsible for execution, and the memory layer runs through everything, providing state support for all stages.


0. Perception Layer (Perception) — The Front Desk of the Restaurant

Role: responsible for receiving customers and understanding all inputs from the external world.

Before taking action, an Agent must first "see" and "hear" external information. Modern Agents are no longer limited to pure text input; they havemultimodal perceptioncapability:

  • Text input: user natural language instructions, document content, and code.
  • Images / Videos: screenshots, design drafts, and charts; Agent can directly "look at pictures" to understand.
  • Structured data: Tables, JSON, database query results.
  • Environment state: In computer-operation agents, the current screen state, webpage DOM structure, etc.
  • Tool return results: The output of the previous tool call serves as new perception input for the next cycle.

The perception layer's inputs are integrated to form the agent's "current context", which is sent to the brain for understanding and decision-making.


1. Brain — That is, the Large Model

Role: The restaurant'shead chef and manager。

This is the most core part of the agent (e.g., GPT-4, Claude, DeepSeek, Qwen).

  • It is responsible forunderstandingwhat you want to eat (understanding intent).
  • It is responsible fordirectingothers to work (decision-making).
  • Without it, the entire restaurant would be paralyzed.

Three Core Things the Brain Does

CapabilityDescriptionRestaurant analogy
Intent understandingParse user input and clarify what the goal isUnderstand what the customer ordered
Reasoning and decision-makingSynthesize context and memory to determine what to do nextThe head chef decides which dish to handle first
Tool call judgmentDetermine whether to call an external tool, which tool to select, and what parameters to passDecide which pot to use and who should buy the ingredients
Key concept:The brain's "intellectual ceiling" determines the upper limit of the entire agent. With the same set of tools and planning framework, connecting a more capable base model often leads to a qualitative leap in task completion quality.

2. Tools — Equipment in the Kitchen

Role:Cookware and helpers。

Having only the head chef (brain) is not enough; you also need pots and pans to cook. For AI agents, tools are execution units that turn decisions into real actions.

Tools can be divided into four categories by purpose:

CategoryCommon toolsFunction
Information retrieval Web search, web scraping, document reading, database queries Obtain real-time or specialized information beyond the agent's own knowledge
Computation and execution Code interpreter, math computation engine, sandbox environment Handle tasks requiring precise computation or program logic
Content generation Image generation, speech synthesis, document export Produce non-textual content
System interaction API interfaces, email, calendar, file operations, message sending Interact with external systems, services, and the real world

Common tool examples:

  • Online search Information retrieval(Like going to the wet market to buy fresh ingredients)
  • Code interpreter. Compute execution(like a precision oven, handling complex computations)
  • Drawing tool Generated Content(Like a plating artist, responsible for aesthetics)
  • API interface System interaction(Like a delivery courier, connecting with the outside world)
Function Calling:presentgenerationbigModelThrough"Function call"MechanismcomeUsageTools。developer预先definitionToolsname ofandParameter description,ModelInInferencewhenwillwithStructuretransform JSON offormOutput"IwantCall哪itemsTools、传什么Parameter", byExternal programnegative责truecorrectExecuteandholdResultreturnGiveModel.

3. Memory — Customer Record Book

Role:The waiter's memory。

You certainly don't like having to repeat every time you go to a restaurant: I don't eat cilantro!

Agent memory is divided into the following types:

  • Short-term memory (In-Context Memory): the context window of the current conversation. It remembers what you just said (e.g., if you just ordered fish, and your next sentence is "lightly spicy", it knows it refers to the fish). Limited by the model's context length, typically between 8K and 200K tokens.
  • Long-term Memory (External Memory): remember your long-term preferences (e.g., you are a vegetarian, or your home address). Usually throughVector database(such as Pinecone, Milvus, Chroma) for persistent storage.
  • Episodic Memory: A record of the execution process of historical tasks, including "how I handled this situation last time," helping the Agent learn from past experience.
  • Semantic Memory: abstract knowledge and facts, usually from content already internalized during the pre-training phase, and can also be dynamically supplemented through RAG (Retrieval-Augmented Generation).

RAG: Giving Agents an "External Knowledge Base"

Retrieval-Augmented Generation (RAG)is currently the most mainstream long-term memory implementation solution. Its core process is as follows:

User question Query Vectorized retrieval. Embedding + Search Retrieve relevant content Top-K Chunks Inject context for generation LLM + Context Answer Answer

4. Planning — Cooking Flow Sheet

Role:The kitchen's meal-serving SOP。

When you order a "Buddha Jumps Over the Wall", the head chef doesn't improvise; instead, they generate a checklist in their mind:

  1. First prepare ingredients (abalone, sea cucumber...)
  2. Then simmer the broth
  3. Finally slow-cook

The same goes for an Agent. When you give it a complex task (e.g., "write a competitive analysis report"), it breaks it down on its own:

  • Step 1: Gather information on competitors A, B, and C.
  • Step 2: Compare their prices and features.
  • Step 3: Write the comparison results into an article.
  • Step 4: Check for typos.

Mainstream Planning Strategies

Planning strategies determine how an Agent "thinks" then "acts"; different strategies differ in reasoning depth and applicable scenarios:

StrategyFull nameCore ideaApplicable scenarios
CoT Chain-of-Thought Write out the reasoning process step by step before giving an answer Mathematical reasoning, logical analysis
ReAct Reasoning + Acting Alternate between "reasoning" and "acting", reasoning again based on the result after each action Dynamic tasks requiring tool calls
ToT Tree-of-Thoughts Explore multiple reasoning branches simultaneously and select the optimal path from them Complex decision-making, creative tasks
Reflection Self-reflection After the task is completed, the Agent critically reviews and corrects its own output Code generation, long-form writing
ReAct example:Agent receives the task "Check tomorrow's Beijing weather and send a reminder" →Thought: Need to check the weather first →Action: Call the weather API →Observation: Returns "It will rain tomorrow" →Thought: The condition holds, need to write a reminder →Action: Call the message-sending tool → Task complete.

5. Agent Run Loop (Agent Loop)

The above components do not exist in isolation; they form a continuously iterativePerception—Thinking—Action—Observationclosed loop, which is the "Agent Loop". The Agent continuously repeats this loop until the task is completed or a termination condition is reached.

Perception Receive input / environment state Thinking LLM reasoning / planning decomposition Action Call tool / Execute operation Observe Get result / Update memory Task complete / Termination condition reached Continue loop (task incomplete)

This loop gives the Agent the ability toself-correct on failureThe ability: if a tool call at some step returns an error or unexpected result, the "Observe" stage feeds this information back to the brain, and the brain adjusts its strategy in the next round of "thinking".


Summary

When you say to the Agent:Please check tomorrow's weather in Beijing; if it's rainy, write a reminder and send it to Xiao Wang.

Here's how the Agent works internally:

  1. Perception layer: Receives the natural language instruction and identifies key entities "Beijing", "tomorrow", "Xiao Wang".
  2. Brain: Hears the instruction and analyzes two conditional tasks: check the weather, and if it rains then send a reminder.
  3. Planning: First check the weather → determine whether it will rain → (if so) write a reminder → send.
  4. Tool: Calls the "weather query tool" and gets the result — it will rain tomorrow.
  5. Memory: Goes to the contacts (memory bank) to look up Xiao Wang's contact information.
  6. Tool: Calls the "send message tool" to send the reminder.
  7. Observe: Confirms the message was sent successfully, the task is complete, and the loop ends.

Runtime process diagram:

The Five Core Components at a Glance

ComponentRestaurant analogyCore responsibilitiesKey technologies
Perception layer Front desk reception Receive multimodal input, build context Multimodal models, OCR, ASR
Brain Head chef and manager Understand intent, reason and make decisions, issue commands LLM、Function Calling
Planning Meal delivery SOP Task decomposition, step sequencing, self-reflection ReAct、CoT、ToT、Reflection
Tool Kitchen utensils and helpers Execute specific operations, connect to the external world Search / Code / API / File system
Memory Customer record book Manage context, store long-term knowledge Vector database, RAG, context window
Other extensions