AI Underlying Architecture
The following diagram shows the complete AI system architecture from foundational large models to agent applications:

| Layer | Core Components | Core Role | Typical Capabilities |
|---|---|---|---|
| Foundation Layer (Model) | LLM & Token, Transformer architecture, Tokenization, attention mechanism | AI's underlying computational core; responsible for text understanding, Token prediction, and language generation | Natural language understanding, text generation, context modeling, probability prediction |
| Context Layer (Memory) | Context Window、Prompt、Memory、RAG | Manages model input context; responsible for short-term memory, long-term memory, and external knowledge injection | Prompt control, session memory, knowledge retrieval, context augmentation |
| Capability Expansion Layer (Tools) | MCP、Tool Calling、API、Database | Makes AI more than just chat; expands real-world operational capabilities through tool calls | Online search, code execution, database queries, API calls |
| Agent Layer (Decision-Making) | Agent、Explore、Plan、Act | AI's autonomous decision-making brain; responsible for goal understanding, task decomposition, planning, and closed-loop execution | Task planning, multi-step reasoning, autonomous decision-making, closed-loop execution |
| Application Layer (Action) | Agent Skill、Workflow、Automation | Targets specific business scenarios; encapsulates Agent capabilities into deployable products | Automated workflows, industry AI assistants, enterprise intelligent systems, AI SaaS applications |
Core Terminology Definitions
- TransformerThe current core architecture of large models, establishing relationships between Tokens through the self-attention mechanism.
- TokenThe smallest unit of text processed by the model; can be a subword, word, character, or symbol.
- PromptThe input instruction given to the model, used to set roles, control behavior, and standardize output format and goals.
- RAGRetrieval-Augmented Generation; first retrieves an external professional knowledge base, then lets the LLM integrate and generate accurate answers.
- MCPModel Context Protocol (Model Context protocol), unifying the communication standard between AI and tools, databases, and external services.
- AgentOn the basis of an LLM, it overlaysgoal + planning + executioncapabilities to form an autonomous intelligent system.
- Tool CallingThe model proactively identifies needs, calls external tools to complete actual operations, and is not limited to pure text responses.
- MemoryProvides AI with short-term conversational memory + long-term persistent memory, solving the problem of context forgetting.
Complete AI Agent Workflow
- user input: The user issues an instruction Prompt, which enters the Context Window.
- Memory Enhancement: The system integrates Memory session memory + RAG external knowledge base to supplement context and domain expertise.
- LLM Inference: Attention computation is performed on Token sequences through the Transformer architecture for semantic understanding and preliminary generation.
- Agent Planning: The agent autonomously judges and decomposes tasks, deciding whether multi-step execution or calling external tools is needed.
- Tool Execution: Based on the MCP protocol, external capabilities such as API, database, and search engine are invoked to complete operations.
- Output Result: Summarizes data returned by tools and model inference results, organizes them and returns to the user, while writing to memory archive.