Agent Context Engineering
Agent 上下文工程(Agent Context Engineering) is a technical practice for systematically designing and optimizing the context information passed to AI Agents. Its goal is to enable Agents to obtain the most effective information within a limited context window, thereby improving the accuracy, efficiency, and reliability of task execution.
Unlike traditional programs, AI Agents do not have fixed execution logic; their behavior is entirely determined by the context they receive—including system prompts, tool descriptions, conversation history, and externally retrieved information.
The core of context engineering is answering a question:In this round of invocation, what does the Agent most need to know?

Context engineering is not one-time prompt writing, but a continuous optimization process that runs through the entire lifecycle of Agent system development, testing, and operations.
Core Elements
Agent context consists of the following layers:
| Layer | Content | Characteristics |
|---|---|---|
| System Prompt | Role definition, behavior rules, output format requirements | Carried on every call, relatively stable |
| Tool Definitions | Names, parameters, and function descriptions of available tools | Consumes a large number of tokens, needs to be concise |
| History | Conversation and tool invocation records of the current session | Continuously grows, requires truncation or summarization strategies |
| Retrieved Context | Content retrieved from external knowledge bases or code repositories | Injected on demand, requires relevance ranking |
| User Input | The user's current instruction or question | Uncontrollable, but can be optimized through clarification |
Core Practices
Context Budget Management
Treat the context window as a finite "budget" and allocate it reasonably across different types of context.
Not all information deserves to occupy context space; every piece of context should have a clear ROI.
上下文预算分配建议(以 200K token 窗口为例): 系统提示 ██████████ 10% (20K token) 工具定义 ████████████████ 20% (40K token) 检索上下文 ████████████████████ 25% (50K token) 历史记录 ████████████████████████ 30% (60K token) 用户输入 ████████ 5%~10% (10-20K token) 预留缓冲 ████ 5% (10K token)
The actual ratio should be dynamically adjusted based on the Agent's specific task type.
If the Agent primarily does code generation, the retrieved context ratio can be increased.
If the Agent is a multi-turn dialogue assistant, the historical record budget should be prioritized.
System Prompt Engineering
The system prompt is the cornerstone of Agent behavior and should follow structured authoring principles.
Authoring principles:
| Principle | Description | Example |
|---|---|---|
| Hierarchical organization | Divide instructions into blocks by role, rules, process, and format | Separate with Markdown headings |
| Positive phrasing | Tell the Agent what to do, not what not to do | "Keep answers concise" is better than "don't be verbose" |
| Provide examples | Use few-shot examples instead of lengthy descriptions | Output format examples are more effective than textual descriptions |
| Clear priorities | Which rule the Agent should follow when rules conflict | "Safety rules take precedence over efficiency requirements" |
| Remove redundancy | Delete rules that are never triggered and duplicate instructions | Regularly review and streamline the system prompt |
The following is an example skeleton of a structured system prompt:
## 角色 你是一名 Python 代码审查助手。 ## 核心规则 - 每个问题附上修复建议 - 按严重程度排序输出 - 不确定时标注「待确认」 ## 工作流程 1. 分析代码变更 2. 按严重程度分类问题 3. 逐一输出问题与建议 ## 输出格式 **问题**:[问题描述] **严重程度**:[高/中/低] **建议**:[修复建议]
Tool Description Optimization
Tool descriptions often occupy the largest share of the context, but many tools are never actually invoked.
Optimization strategies include:
| Strategy | Method | Effect |
|---|---|---|
| Tool reduction | Remove tool definitions unrelated to the current task | Reduce tool context by 30%~50% |
| Concise descriptions | Describe each tool's purpose and parameters in a single sentence | Improve the model's accuracy in understanding tools |
| Parameter constraints | Clarify parameter usage scenarios and limitations in the description | Reduce parameter errors in tool calls |
| Grouped registration | Dynamically register different tool sets by task stage | Avoid "choice difficulty" caused by too many tools |
Any non-zero tool definition is a burden. Every retained tool requires a trade-off between description quality and token cost.
History Compression Strategies
In multi-turn conversations, the history grows quickly and consumes context.
It is necessary to choose an appropriate compression strategy to balance context completeness and window limits.
常见历史压缩策略对比: 策略 方法 适用场景 ───────────────────────────────────────────────── 滑动窗口 保留最近 N 轮完整记录 短对话、实时交互 阶梯摘要 每轮做一次增量摘要 长对话、客服场景 分层摘要 近期保留完整、远期做摘要 文档生成、复杂任务 关键轮标记 标记重要轮次,其余丢弃 调试、分步执行任务
Hierarchical summarization is the most commonly used strategy: keep the most recent 3-5 turns complete, and replace earlier conversations with structured summaries.
A structured summary should contain the key information of the original conversation: what the user's goal was, what the Agent did, what output it produced, and what errors it encountered.
Quality Control for Retrieved Context
When an Agent relies on an external knowledge base, the quality of the retrieved context directly determines the quality of the output.
Key points for improving the quality of retrieved context:
| Stage | Common issues | Improvement methods |
|---|---|---|
| Query rewriting | User input is not precise enough | Have the Agent rewrite the query before retrieving |
| Relevance filtering | Retrieved results contain irrelevant content | Set a relevance threshold, and proactively alert when it is insufficient |
| Source annotation | The Agent cannot judge the credibility of information | Attach the source path and update time |
| Length trimming | Retrieved results occupy too much context | Truncate by paragraph, keeping the most relevant segments |
Common Patterns
Progressive Disclosure
Instead of stuffing all information into the context at once, disclose it progressively according to the execution stage.
The initial context only contains the system prompt and the tools necessary for the current stage. After the Agent completes the current stage, the tools and rules for the next stage are injected.
This mode keeps the "density" of the context within a reasonable range, preventing the Agent from being disturbed by irrelevant information.
Context Compression Chain
When a task requires processing large amounts of data, split it into multiple steps using chained calls, keeping only key intermediate results at each step.
步骤1:读取全部日志文件 → 输出 { 错误数量,时间范围,错误类型列表 }
步骤2:基于步骤1的输出,分析 Top 3 错误 → 输出 { 根因分析,修复建议 }
步骤3:基于步骤2的输出,生成修复 PR → 输出 { PR 标题,描述,代码变更 }
The input of each link is the refined summary from the previous step, rather than the raw data, keeping the context at a controllable size at all times.
Context Watermarking
Place "watermark" markers at key positions in the context to help monitor and diagnose Agent behavior.
上下文水印示例: <system-reminder> 当前使用的上下文策略版本:v2.3 检索到的文档来源:docs/api-reference/ 共 3 个文件 已消耗上下文:85,000 / 200,000 tokens </system-reminder>
These watermarks do not directly affect Agent behavior, but during debugging they can quickly locate whether context injection was executed as expected.
Evaluation and Iteration
The effectiveness of context engineering needs to be verified through evaluation and cannot be judged by intuition alone.
Commonly used evaluation dimensions:
| Dimension | Evaluation method | Key metrics |
|---|---|---|
| Task completion rate | Run the Agent on a standard test set | Success rate, failure reason distribution |
| Context efficiency | Count the context usage of each call | Average token consumption, context utilization rate |
| Tool call quality | Analyze tool call accuracy and parameter correctness rate | Invalid call ratio, retry count |
| Response consistency | Multiple tests with the same input | Output stability, format compliance rate |
Context engineering is not a one-time configuration; it requires establishing an iterative loop of "test → analyze → optimize → retest."
Notes
A larger context window is not always better. Excessive context can dilute the model's attention on key information and actually reduce output quality.
Blindly adding rules can make system prompts bloated and self-contradictory. Each new rule needs to be reviewed for conflicts with existing rules.
Changes to tool definitions may affect the behavior of existing agents. After modification, regression testing should be performed in typical scenarios.
The history compression strategy needs to be chosen based on the agent's task characteristics; no single strategy can perform optimally in all scenarios.
Sensitive information in the context (API keys, internal paths, etc.) can propagate through tool calls and logs, so proper desensitization is needed.
Related Concepts
| Concept | Description |
|---|---|
| Prompt Engineering | Focus on prompt design and optimization for a single call |
| Context Window | The maximum number of tokens a model can process at once |
| Token | The smallest semantic unit of text processed by a model |
| RAG (Retrieval-Augmented Generation) | An architectural pattern that expands context through external retrieval |
| Few-shot Prompting | Provide examples in the context to guide model output |
| Chain-of-Thought | Let the model show its reasoning process in the context |