Agent Architecture

An Agent (intelligent agent) refers to an AI system that can autonomously perceive its environment, reason, and take actions to achieve goals.

From the simplest loop to multi-agent collaboration, this article gives you a thorough understanding of the principles, diagrams, and applicable scenarios of six mainstream architectures, helping you make appropriate technical choices in real projects.


What is Agent Architecture

In AI application development,Agent (intelligent agent)Refers to an AI system that can perceive its environment, make autonomous decisions, and take actions.

Unlike traditional "question-answer" style large model calls, an Agent can continuously perform multi-step operations, call tools, and even coordinate other Agents to complete complex tasks.

Agent ArchitectureRefers to the organization of components in an Agent system, which determines the capability boundaries, reliability, flexibility, and applicable scenarios of the Agent.

The essence of Agent: Perception → Reasoning → Action
Agent Core Loop Shows the three core stages of an Agent: perceiving the environment, reasoning and deciding, executing actions, forming a loop. Perceive environment Reason and decide Execute action Observe results, continue the loop

The way an Agent works is essentially aLoop— perceive the current state, reason about the next step, execute an action, perceive again... until the task is complete.

The difference between architectures lies in how to organize and extend this basic loop.

This article assumes you are already familiar with the basic concepts of large language models (LLMs) and the basic idea of tool calling (Function Calling). If you are not familiar yet, you can learn these two fundamentals first.


Architecture 1: Single Agent Loop

Best for beginners Simple to implement

The most basic and intuitive architecture: one Agent independently completes all tasks from start to finish.

The single Agent loop directly embodies theReAct pattern(Reasoning + Acting): every step is "think first, then act." The LLM serves as the brain, and tool calling is its hands.

Figure 1 — Single Agent Loop flow
Single Agent Loop architecture diagram Illustrates the complete working loop of a single Agent: receive task, perceive environment, reason and decide, call tools, then judge based on the result whether to continue looping or exit when the task is complete. User task input Perceive · Gather context Read files, get status Reason · Decide next step Analyze situation, select tool Act · Call tool Bash / Read / Edit... Done Output result Continue loop Agent loop

How it works

Perceive:Read the current state — file contents, environment variables, output from previous steps — and integrate it into the current context.

Reason:The LLM decides the next action based on the context — which tool to call, what parameters to pass — or determines whether the task is complete.

Act:Execute tool calls, such as reading/writing files or searching the web. The tool's execution results are appended to the context, and it proceeds to the next loop iteration.

Every tool call result is written back to the context window. Therefore, as the task progresses, the context keeps growing until it reaches the LLM's context window limit — this is the primary bottleneck of the single Agent loop.

Example

# Simplified implementation of the single Agent loop — showing the core logic of the ReAct pattern

class SimpleAgent:
    """Basic structure of the single Agent loop"""

    def __init__(self, model, tools, max_turns=10):
        self.model = model          # Large language model
        self.tools = tools          # Available tool list
        self.max_turns = max_turns  # Maximum loop iterations, preventing infinite loop

    def run(self, task: str) -> str:
        """Main loop for executing the task"""
        context = f"User task: {task}"

        for turn in range(self.max_turns):
            # Step 1: Think — let the model decide the next step
            response = self.model.think(context)

            # If the model believes the task is complete, return the final answer
            if response.is_final():
                return response.content

            # Step 2: Act — call the tool chosen by the model
            tool_name = response.tool_name
            tool_args = response.tool_args
            tool_result = self.tools[tool_name](**tool_args)

            # Step 3: Feed the tool result back to the model and proceed to the next round
            context += f"\nTool {tool_name} returned: {tool_result}"

        return "Maximum rounds reached, task not completed"

# Usage example
agent = SimpleAgent(model=llm, tools={
    "read_file": read_file,
    "search_code": search_code,
    "run_test": run_test
})
result = agent.run("Fix the type error in user.py in the example project")

Advantages

  • Simplest to implement, easy to debug
  • Suitable for scenarios with clear task boundaries
  • Supported by almost all Agent frameworks

Disadvantages

  • Context window easily fills up
  • Prone to "going off course" in complex tasks
  • Cannot process multiple subtasks in parallel

Best use cases:Fix a bug, write a function, answer a specific question. The task is clear, with moderate complexity, and does not require parallelism or multi-role collaboration.

If your task is expected to require more than 15 rounds of tool calls, a single Agent loop may not be the best choice. Consider using multi-Agent collaboration or a plan-and-execute architecture to break down the complexity.


Architecture 2: Plan & Execute

Intuitive Reviewable

Separates "figuring out what to do" and "actually doing it" into two independent phases, improving task predictability and auditability.

The Plan + Execute architecture splits the Agent's work into two distinct phases: firstPlan, thenExecute。

In the planning phase, the model does not perform any actions; it only generates a detailed list of execution steps. In the execution phase, the system completes each step in order. This separation allows users to review the plan before execution, similar to Claude Code's Plan Mode.

Figure 2 — Plan + Execute architecture (with dynamic replanning)
Plan + Execute architecture diagram Shows the two-phase process: the planning phase generates a list of execution steps, the execution phase completes each step in sequence, and can dynamically replan based on results Planning phase User task Planner Generate a step-by-step execution plan Execution plan Step 1 Gather information Step 2 Analysis processing Step 3 Output result Execution phase Executor Complete each step in order Complete the task Dynamic re-planning (Optional)

Two variants

VariantBehaviorTypical scenario
Static planningThe plan is generated all at once and executed linearly in order without mid-course adjustments.Tasks with fixed processes and well-defined steps, such as data migration scripts.
Dynamic planningRe-evaluate after each step and adjust the subsequent plan based on the results.Tasks with uncertain outcomes, such as debugging and exploratory data analysis.

Dynamic planning is more robust, but it is more complex to implement, and re-planning at each step consumes extra tokens.

Claude Code's Plan Mode embodies this architecture—after clicking "Plan", the AI first outputs a detailed plan for your review, and only begins execution after confirmation. This greatly enhances the user's sense of control.

Example

# Simplified implementation of the Plan & Execute architecture

class PlanExecuteAgent:
    """Agent that plans first, then executes"""

    def plan(self, task: str) -> list:
        """Phase 1: Generate an execution plan"""
        plan = self.model.generate(f"""
Please break down the following task into a list of executable steps:
Task: {task}
Return a JSON-formatted list of steps, each containing:
- step_id: step number
- description: step description
- tool: name of the tool to call
        """
)
        return plan

    def execute(self, plan: list, dynamic: bool = False) -> str:
        """Phase 2: Execute the plan step by step"""
        results = []
        remaining_plan = plan.copy()

        while remaining_plan:
            step = remaining_plan.pop(0)
            output = self.tools[step["tool"]](step["description"])
            results.append({"step": step["step_id"], "output": output})

            if dynamic and remaining_plan:
                # Dynamic planning: re-evaluate the subsequent plan based on current results
                remaining_plan = self.replan(remaining_plan, results)

        return self.summarize(results)

# Usage example
agent = PlanExecuteAgent()
plan = agent.plan("Add user authentication functionality to the example project")
# Humans can review the plan first and execute after confirming it is reasonable
result = agent.execute(plan, dynamic=True)

Advantages

  • The plan can be manually reviewed before execution.
  • Clear separation of reasoning and execution responsibilities.
  • Friendly for long tasks.

Disadvantages

  • The initial plan may not be accurate enough.
  • The two-phase approach adds latency.
  • The static version struggles to handle unexpected situations.

The cost of Plan & Execute is increased reasoning rounds, which is wasteful for simple tasks. If a task can be completed in 3 steps, using a single-agent loop directly is more efficient.


Architecture 3: Multi-Agent Collaboration

Production recommendation Complex tasks

An Orchestrator is responsible for task decomposition and scheduling; multiple Subagents each perform their own duties, completing subtasks in parallel or sequentially, and results are gathered back to the Orchestrator for synthesis.

When a single Agent faces an insufficient context window or overly complex tasks, the multi-Agent architecture offers an elegant solution:Have multiple specialized sub-Agents work in parallel, while an Orchestrator coordinates the overall situation.

Figure 3 — Multi-Agent Collaboration Architecture (Orchestrator + Subagents)
Multi-Agent collaboration architecture diagram The Orchestrator distributes tasks to three specialized sub-Agents, which execute in parallel and then return the results to the coordinator Orchestrator Task decomposition · scheduling · result synthesis Distribute tasks Execute in parallel Subagent A Code review Subagent B Security detection Subagent C Performance analysis Return results Result aggregation Synthesize outputs from all sub-Agents

Independent context is a core advantage

Each sub-Agent has anindependent context window. The code review Agent's deep reading of auth.py does not affect the performance analysis Agent's judgment; the large amount of intermediate output from the security detection Agent does not crowd out other Agents' space.

A Subagent is transient and isolated—it is destroyed after completing a task. Agent Teams, on the other hand, are multiple independent Agent instances collaborating over a long period and messaging each other, more like a real team.

Example

# Simplified implementation of multi-Agent collaboration

class Orchestrator:
    """Orchestrator: responsible for task decomposition, distribution, and result aggregation"""

    def __init__(self):
        self.subagents = {
            "code_review": Subagent(
                name="Code review",
                tools=["read_file", "static_analysis"],
                system_prompt="You are a code review expert..."
            ),
            "security": Subagent(
                name="Security detection",
                tools=["scan_vulnerability", "check_deps"],
                system_prompt="You are a security detection expert..."
            ),
            "performance": Subagent(
                name="Performance analysis",
                tools=["profile_code", "analyze_complexity"],
                system_prompt="You are a performance analysis expert..."
            )
        }

    def handle_task(self, task: str) -> dict:
        # Step 1: Analyze the task and decide which Subagents are needed
        needed = self.plan(task)

        # Step 2: Distribute in parallel (each Subagent works simultaneously with independent context)
        results = {}
        for agent_name in needed:
            sub_task = self.decompose(task, agent_name)
            results[agent_name] = self.subagents[agent_name].run(sub_task)

        # Step 3: Aggregate the results of each Subagent and produce a comprehensive output
        return self.synthesize(task, results)

# Usage Example: One Run, Three-Dimensional Parallel Analysis
orch = Orchestrator()
report = orch.handle_task(Review PR #42 of the example project)

Advantages

  • Naturally supports parallelism, fast
  • Sub-agents are independent, contexts do not interfere with each other
  • Can specialize the role of each sub-agent

Disadvantages

  • Complex coordination logic, difficult to debug
  • Higher token cost for parallel multi-agent execution
  • The Orchestrator itself may become a bottleneck

The main cost of multi-agent collaboration is orchestration overhead. If subtasks are very simple (each requiring only 1-2 steps), orchestration overhead may exceed the cost of the actual work, in which case a single agent is more appropriate.


Architecture 4: Reflection and Self-Correction

High-quality output Easy to integrate

Add quality assessment to the agent's output stage; if unsatisfactory, regenerate or correct, forming an internal iteration loop.

The reflection architecture adds a "quality inspection step" for the agent: after each output generation, aCriticevaluates the quality, and if it does not meet the standard, requests corrections until the output satisfies the criteria.

This is like a developer running tests on their own code after writing it—self-checking before delivery.

Figure 4 — Reflection and Self-Correction Architecture
Reflection and Self-Correction Architecture Diagram After the agent executes the task, the Critic scores the output quality; if it passes, the task is complete; otherwise, it enters the correction stage. Executor Generates initial output Critic Evaluates quality and correctness Pass Output final result Fail Reviser Improves output based on feedback Regenerate (Usually with a maximum iteration limit set)

Two implementation approaches

ApproachMechanismAdvantagesDisadvantages
Self-reflectionThe same model executes first, then evaluates its own outputSimple implementation, no additional model costThe model may be "blind" to its own errors
Critic modelUse an independent critic model to evaluate the executor model's outputMore objective, can detect blind spots of the executor modelAdds model call cost and latency

"Write unit tests → Run tests → Observe failures → Fix code → Run again" — this is a classic application of the reflection architecture. The test results themselves serve as the Critic's feedback signal.

Example

# Simplified implementation of the reflection architecture

class ReflectiveAgent:
    """Agent with self-reflection capability"""

    def __init__(self, model, tools, max_reflections=3):
        self.model = model
        self.tools = tools
        self.max_reflections = max_reflections  # Maximum number of corrections to prevent infinite loops

    def run(self, task: str) -> str:
        # Step 1: Execute normally, produce initial output
        output = self.model.generate(task)

        for i in range(self.max_reflections):
            # Step 2: Reflection — evaluate output quality
            critique = self.model.generate(f"""
Please strictly evaluate the following output:
Original task: {task}
Current output: {output}
Check: factual errors? logic gaps? missing information? formatting issues?
If the output is flawless, reply "PASS".
            """
)

            if "PASS" in critique:
                break  # Output passed review

            # Step 3: Correction — improve based on critique
            output = self.model.generate(f"""
Original task: {task}
Previous output: {output}
Issue feedback: {critique}
Please correct the output based on the feedback.
            """
)

        return output

# Usage example
agent = ReflectiveAgent(model=llm, tools={})
code = agent.run("Write a Python function that implements AES encryption for the EXAMPLE string")
# After generating the code, the Agent self-checks the encryption implementation and key handling,
# automatically fixes vulnerabilities after detection, ensuring secure and reliable output

Advantages

  • Significantly improves output quality
  • Allows setting clear quality standards
  • Suitable for tasks with objective evaluation criteria

Disadvantages

  • Multiple iterations increase latency and cost
  • Need to set a maximum iteration count to prevent infinite loops
  • Limited effectiveness when evaluation criteria are difficult to formalize

Each iteration of the reflection loop is an additional LLM call, which significantly increases latency. In addition, an upper limit on the number of reflections needs to be set, otherwise the model may fall into an "never satisfied" infinite loop.


Architecture 5: RAG + Agent (Retrieval-Augmented Agent)

Knowledge-intensive tasks Large knowledge bases

Add vector retrieval capability to the Agent's toolset, allowing the Agent to dynamically query external knowledge bases during reasoning, overcoming the limitations of the context window.

RAG (Retrieval-Augmented Generation) is originally a technique that allows LLMs to query external knowledge bases. When combined with an Agent, it becomes more powerful: the Agent canproactively decide when to retrieve and what to retrieve, rather than passively retrieving once each time.

Figure 5 — RAG + Agent retrieval-augmented architecture
RAG + Agent Architecture Diagram During reasoning and decision-making, the Agent actively initiates retrieval from external knowledge bases, incorporates the retrieved relevant content into context, and then performs more accurate reasoning and actions. User Query Active Retrieval Decide when and what to retrieve Knowledge Base Documents / Vector Database Codebase / Memory Query Return Enhanced Context Enhanced Reasoning Reasoning Combined with Knowledge Base Content On-demand Re-retrieval Execute Actions Precise Knowledge-Based Operations

Key differences from ordinary RAG

Ordinary RAG is "passive and one-shot": when a user asks a question, it retrieves once and stuffs the results into the Prompt.

RAG + Agent is differentThe Agent autonomously determines which stage of reasoning needs additional knowledge and what to retrieve, and can query the knowledge base multiple times until it obtains enough information to complete the task.

When the codebase Q&A Agent is asked "Why is the singleton pattern used here?", it proactively retrieves project documentation, design decision records, and relevant code files, rather than guessing only from the LLM's training memory.

Example

# Simplified Implementation of RAG + Agent

class RAGAgent:
    """Agent with dynamic retrieval capability"""

    def __init__(self, model, vector_db, max_retrievals=5):
        self.model = model
        self.vector_db = vector_db  # Vector Database
        self.max_retrievals = max_retrievals

    def should_retrieve(self, context: str, question: str) -> bool:
        """Agent decides whether it needs to retrieve more information"""
        decision = self.model.generate(f"""
Current known information: {context}
Current question: {question}
Is the existing information sufficient to answer the question? Answer YES or NO.
        """
)
        return "NO" in decision

    def run(self, task: str) -> str:
        context = ""
        retrieval_count = 0

        while retrieval_count < self.max_retrievals:
            # Agent autonomously determines whether retrieval is needed
            if not self.should_retrieve(context, task):
                break

            # Agent autonomously decides what to retrieve
            search_query = self.model.generate(f"""
Task: {task}
Existing information: {context}
To complete the task, what information should be retrieved next?
            """
)

            # Perform retrieval, append results to context
            docs = self.vector_db.search(search_query)
            context += "\n".join(docs)
            retrieval_count += 1

        # Synthesize all information to generate the final answer
        return self.model.generate(f"Task: {task}\nReference materials: {context}")

# Usage Example
agent = RAGAgent(model=llm, vector_db=example_docs_db)
answer = agent.run("How do you configure a database connection pool in the EXAMPLE framework?")
# Agent first retrieves "connection pool configuration" and finds mentions of "maximum connections"
# If it doesn't understand, it retrieves "maximum connections best practices" again
# Finally, it synthesizes multi-round retrieval results to provide a complete answer

Advantages

  • Break through context window limitations
  • Output is well-documented, reducing hallucinations
  • Knowledge base can be updated independently

Disadvantages

  • Retrieval quality affects overall results
  • Maintenance cost of vector databases
  • Retrieval latency increases response time

Architecture 6: Workflow Orchestration (Workflow / DAG)

Preferred for production High reliability

Solidify Agent behavior into a directed acyclic graph (DAG), where each node is an LLM call or tool call, edges represent data dependencies, and execution is driven by the framework.

This is an Agent architecture closest to traditional software engineering. The biggest difference from the previous architectures is:The Agent's autonomous decision-making space is limited to within a single node, and the flow between nodes is predefined and unchangeable.

Figure 6 — Workflow DAG Architecture (Directed Acyclic Graph)
Workflow DAG architecture diagram Directed acyclic graph illustration: After data input, Task A and Task B execute in parallel; after both complete, the aggregation node processes; finally, output the result. Data input Task A Independent parallel execution Task B Independent parallel execution Parallelizable Triggered only after both A and B complete Aggregation node Wait for dependencies, integrate results Output

Pure Agent vs DAG Workflow

FeaturePure AgentDAG Workflow
Flow controlModel autonomously decides the next stepDetermined by predefined DAG graph
PredictabilityLow, execution path may differ each timeHigh, execution path is fixed
DebuggabilityDifficult, relies on log tracingEasy, each node's input and output are clear
Fault toleranceRelies on model self-recoveryFramework provides retry and resume from checkpoint
FlexibilityHigh, can handle unexpected situationsLow, can only follow predefined paths

The "Acyclic" property of DAG means the workflow isDeterministic— no infinite loops, execution path can be fully predicted, and failed nodes can be retried individually.

DAG is the extreme of "low autonomy, high predictability." You need to pre-design the entire process. This is not a disadvantage, but a deliberate design trade-off — in production environments, determinism is sometimes more important than flexibility.

Example

# Simplified definition of DAG workflow (LangGraph style)

from langgraph import StateGraph

# Define workflow state — data object passed between nodes
class PipelineState:
    raw_data: str = ""         # Raw input data
    cleaned_data: str = ""     # Cleaned data
    analysis_result: dict = {} # Analysis result
    final_report: str = ""     # Final report

# Define DAG nodes — each node is an independent processing unit
def extract_data(state: PipelineState) -> PipelineState:
    """Node 1: Extract raw data from example database"""
    state.raw_data = query_database("SELECT * FROM logs")
    return state

def clean_data(state: PipelineState) -> PipelineState:
    """Node 2: Clean data (deduplicate, standardize format)"""
    state.cleaned_data = preprocess(state.raw_data)
    return state

def analyze_data(state: PipelineState) -> PipelineState:
    """Node 3: Statistical analysis"""
    state.analysis_result = statistical_analysis(state.cleaned_data)
    return state

def generate_report(state: PipelineState) -> PipelineState:
    """Node 4: Generate report using LLM"""
    state.final_report = llm.generate(
        f"Generate report based on the following analysis results: {state.analysis_result}"
    )
    return state

# Build DAG: define nodes and edges (data flow)
workflow = StateGraph(PipelineState)
workflow.add_node("extract", extract_data)
workflow.add_node("clean", clean_data)
workflow.add_node("analyze", analyze_data)
workflow.add_node("report", generate_report)

# Define edges: extract → clean → analyze → report
workflow.add_edge("extract", "clean")
workflow.add_edge("clean", "analyze")
workflow.add_edge("analyze", "report")
workflow.set_entry_point("extract")
workflow.set_finish_point("report")

# Compile and run
app = workflow.compile()
result = app.invoke(PipelineState())
print(result.final_report)

Advantages

  • Predictable, auditable, retryable
  • Supports parallel acceleration
  • High engineering maturity, ops-friendly

Disadvantages

  • Flow requires pre-design, low flexibility
  • Difficult to handle unexpected situations
  • Requires learning orchestration frameworks

Horizontal Comparison and How to Choose

Compare six architectures across multiple dimensions to help you quickly identify the right option.

ArchitectureAutonomyPredictabilityParallel capabilitySuitable task complexityTypical implementation
Single Agent loopHighLowNoneMedium-lowClaude Code default mode
Planning + ExecutionMediumMediumPartialMedium-highClaude Code Plan Mode
Multi-Agent collaborationHighLowstrongHighAutoGen, CrewAI
Reflection / self-correctionMediumMediumNoneMediumReflexion, Self-Refine
RAG + AgentHighMediumNoneMedium-highLangChain RAG Agent
Workflow orchestrationLowHighstrongHigh (fixed flow)LangGraph, Prefect

Common combination patterns

Real production systems often combine multiple architectures. Here are several mature combination patterns:

Workflow orchestration + Multi-Agent: Use DAG to define the main flow, with each node being an independent Agent. For example, in a CI/CD pipeline, the code review node is a Code Review Agent, and the security scan node is a Security Agent.

Multi-Agent + RAG: Multiple Subagents share the same vector knowledge base, each retrieving on demand based on its subtask. For example, in a customer service system, both the order query Agent and the refund processing Agent query the same knowledge base but retrieve different content.

Planning execution + Reflection: Plan first, then execute, but add a reflection step after each step to ensure quality. Suitable for tasks with extremely high quality requirements.

If you are unsure where to start, start with a single-Agent loop. It is the easiest to implement and debug. When you find the context window is no longer sufficient, consider multi-Agent; when you need quality assurance, add a reflection layer; when the process stabilizes, refactor it into a DAG to improve reliability.

Common misconceptions

Misconception 1: The more complex the architecture, the better

Multi-Agent collaboration looks powerful, but if your task can be completed in 5 steps with a single-Agent loop, introducing orchestration overhead actually reduces efficiency. Principle: use the simplest architecture that meets the requirements.

Misconception 2: Reflection will definitely improve quality

The effectiveness of self-reflection depends on the model's self-evaluation ability. If quality requirements are extremely strict, consider using a Critic model or introducing external validation (such as automated code testing).

Misconception 3: DAG workflows do not need Agents

DAG defines the process skeleton, but each node can still be an Agent call internally. Workflow orchestration and Agent capabilities are not mutually exclusive but complementary — DAG provides reliability, and Agents provide flexibility.

Misconception 4: A larger context window means RAG is not needed

Even if the model supports a 1M token context window, stuffing all documents into it is still not optimal. The value of RAG is not just "being able to fit it all," but alsoprecise retrieval— reducing noise, lowering inference costs, and improving answer accuracy.


Summary

Six Agent architectures cover the full spectrum from highly autonomous to highly controllable.

When choosing an architecture, there are only two core considerations:How much flexibility you need to handle unexpected situations, andHow much certainty you need to ensure reliable results。

Start simple and add complexity only when truly needed — this is the first principle of Agent architecture selection.

Other extensions