Agent Context Engineering

Agent 上下文工程(Agent Context Engineering) is a technical practice for systematically designing and optimizing the context information passed to AI Agents. Its goal is to enable Agents to obtain the most effective information within a limited context window, thereby improving the accuracy, efficiency, and reliability of task execution.

Unlike traditional programs, AI Agents do not have fixed execution logic; their behavior is entirely determined by the context they receive—including system prompts, tool descriptions, conversation history, and externally retrieved information.

The core of context engineering is answering a question:In this round of invocation, what does the Agent most need to know?

Context engineering is not one-time prompt writing, but a continuous optimization process that runs through the entire lifecycle of Agent system development, testing, and operations.

Core Elements

Agent context consists of the following layers:

LayerContentCharacteristics
System PromptRole definition, behavior rules, output format requirementsCarried on every call, relatively stable
Tool DefinitionsNames, parameters, and function descriptions of available toolsConsumes a large number of tokens, needs to be concise
HistoryConversation and tool invocation records of the current sessionContinuously grows, requires truncation or summarization strategies
Retrieved ContextContent retrieved from external knowledge bases or code repositoriesInjected on demand, requires relevance ranking
User InputThe user's current instruction or questionUncontrollable, but can be optimized through clarification

Core Practices

Context Budget Management

Treat the context window as a finite "budget" and allocate it reasonably across different types of context.

Not all information deserves to occupy context space; every piece of context should have a clear ROI.

上下文预算分配建议(以 200K token 窗口为例):

系统提示  ██████████ 10% (20K token)
工具定义  ████████████████ 20% (40K token)
检索上下文 ████████████████████ 25% (50K token)
历史记录  ████████████████████████ 30% (60K token)
用户输入  ████████ 5%~10% (10-20K token)
预留缓冲  ████ 5% (10K token)

The actual ratio should be dynamically adjusted based on the Agent's specific task type.

If the Agent primarily does code generation, the retrieved context ratio can be increased.

If the Agent is a multi-turn dialogue assistant, the historical record budget should be prioritized.

System Prompt Engineering

The system prompt is the cornerstone of Agent behavior and should follow structured authoring principles.

Authoring principles:

PrincipleDescriptionExample
Hierarchical organizationDivide instructions into blocks by role, rules, process, and formatSeparate with Markdown headings
Positive phrasingTell the Agent what to do, not what not to do"Keep answers concise" is better than "don't be verbose"
Provide examplesUse few-shot examples instead of lengthy descriptionsOutput format examples are more effective than textual descriptions
Clear prioritiesWhich rule the Agent should follow when rules conflict"Safety rules take precedence over efficiency requirements"
Remove redundancyDelete rules that are never triggered and duplicate instructionsRegularly review and streamline the system prompt

The following is an example skeleton of a structured system prompt:

## 角色
你是一名 Python 代码审查助手。

## 核心规则
- 每个问题附上修复建议
- 按严重程度排序输出
- 不确定时标注「待确认」

## 工作流程
1. 分析代码变更
2. 按严重程度分类问题
3. 逐一输出问题与建议

## 输出格式
**问题**:[问题描述]
**严重程度**:[高/中/低]
**建议**:[修复建议]

Tool Description Optimization

Tool descriptions often occupy the largest share of the context, but many tools are never actually invoked.

Optimization strategies include:

StrategyMethodEffect
Tool reductionRemove tool definitions unrelated to the current taskReduce tool context by 30%~50%
Concise descriptionsDescribe each tool's purpose and parameters in a single sentenceImprove the model's accuracy in understanding tools
Parameter constraintsClarify parameter usage scenarios and limitations in the descriptionReduce parameter errors in tool calls
Grouped registrationDynamically register different tool sets by task stageAvoid "choice difficulty" caused by too many tools

Any non-zero tool definition is a burden. Every retained tool requires a trade-off between description quality and token cost.

History Compression Strategies

In multi-turn conversations, the history grows quickly and consumes context.

It is necessary to choose an appropriate compression strategy to balance context completeness and window limits.

常见历史压缩策略对比:

策略     方法            适用场景
─────────────────────────────────────────────────
滑动窗口   保留最近 N 轮完整记录    短对话、实时交互
阶梯摘要   每轮做一次增量摘要     长对话、客服场景
分层摘要   近期保留完整、远期做摘要  文档生成、复杂任务
关键轮标记  标记重要轮次,其余丢弃   调试、分步执行任务

Hierarchical summarization is the most commonly used strategy: keep the most recent 3-5 turns complete, and replace earlier conversations with structured summaries.

A structured summary should contain the key information of the original conversation: what the user's goal was, what the Agent did, what output it produced, and what errors it encountered.

Quality Control for Retrieved Context

When an Agent relies on an external knowledge base, the quality of the retrieved context directly determines the quality of the output.

Key points for improving the quality of retrieved context:

StageCommon issuesImprovement methods
Query rewritingUser input is not precise enoughHave the Agent rewrite the query before retrieving
Relevance filteringRetrieved results contain irrelevant contentSet a relevance threshold, and proactively alert when it is insufficient
Source annotationThe Agent cannot judge the credibility of informationAttach the source path and update time
Length trimmingRetrieved results occupy too much contextTruncate by paragraph, keeping the most relevant segments

Common Patterns

Progressive Disclosure

Instead of stuffing all information into the context at once, disclose it progressively according to the execution stage.

The initial context only contains the system prompt and the tools necessary for the current stage. After the Agent completes the current stage, the tools and rules for the next stage are injected.

This mode keeps the "density" of the context within a reasonable range, preventing the Agent from being disturbed by irrelevant information.

Context Compression Chain

When a task requires processing large amounts of data, split it into multiple steps using chained calls, keeping only key intermediate results at each step.

步骤1:读取全部日志文件 → 输出 { 错误数量,时间范围,错误类型列表 }

步骤2:基于步骤1的输出,分析 Top 3 错误 → 输出 { 根因分析,修复建议 }

步骤3:基于步骤2的输出,生成修复 PR → 输出 { PR 标题,描述,代码变更 }

The input of each link is the refined summary from the previous step, rather than the raw data, keeping the context at a controllable size at all times.

Context Watermarking

Place "watermark" markers at key positions in the context to help monitor and diagnose Agent behavior.

上下文水印示例:

<system-reminder>
当前使用的上下文策略版本:v2.3
检索到的文档来源:docs/api-reference/ 共 3 个文件
已消耗上下文:85,000 / 200,000 tokens
</system-reminder>

These watermarks do not directly affect Agent behavior, but during debugging they can quickly locate whether context injection was executed as expected.


Evaluation and Iteration

The effectiveness of context engineering needs to be verified through evaluation and cannot be judged by intuition alone.

Commonly used evaluation dimensions:

DimensionEvaluation methodKey metrics
Task completion rateRun the Agent on a standard test setSuccess rate, failure reason distribution
Context efficiencyCount the context usage of each callAverage token consumption, context utilization rate
Tool call qualityAnalyze tool call accuracy and parameter correctness rateInvalid call ratio, retry count
Response consistencyMultiple tests with the same inputOutput stability, format compliance rate

Context engineering is not a one-time configuration; it requires establishing an iterative loop of "test → analyze → optimize → retest."


Notes

A larger context window is not always better. Excessive context can dilute the model's attention on key information and actually reduce output quality.

Blindly adding rules can make system prompts bloated and self-contradictory. Each new rule needs to be reviewed for conflicts with existing rules.

Changes to tool definitions may affect the behavior of existing agents. After modification, regression testing should be performed in typical scenarios.

The history compression strategy needs to be chosen based on the agent's task characteristics; no single strategy can perform optimally in all scenarios.

Sensitive information in the context (API keys, internal paths, etc.) can propagate through tool calls and logs, so proper desensitization is needed.


Related Concepts

ConceptDescription
Prompt EngineeringFocus on prompt design and optimization for a single call
Context WindowThe maximum number of tokens a model can process at once
TokenThe smallest semantic unit of text processed by a model
RAG (Retrieval-Augmented Generation)An architectural pattern that expands context through external retrieval
Few-shot PromptingProvide examples in the context to guide model output
Chain-of-ThoughtLet the model show its reasoning process in the context
Other extensions