LangChain Middleware
LangChain Middleware is LangChain's most powerful feature. It allows you to insert custom logic at various stages of Agent execution to implement retries, degradation, caching, content filtering, logging, and other functions — without modifying the Agent's own code.
What is Middleware
Middleware consists ofhooks (Hook)in the Agent execution flow. Each hook lets you execute custom code at specific points in time:
Example
# Assume the Agent execution flow is like this:
# 1. User input → 2. Model thinking → 3. May call tools → 4. Model thinks again → 5. Output result
# Middleware lets you insert custom logic between these 5 stages:
# 1. User input
# ↓ [before_agent hook: logging, permission checks]
# 2. Model thinking
# ↓ [before_model hook: message preprocessing]
# ↓ [wrap_model_call hook: retry, degradation, caching]
# ↓ [after_model hook: content moderation]
# 3. Tool execution
# ↓ [wrap_tool_call hook: tool call retry]
# 4. Back to model thinking (loop until complete)
# ↓ [after_agent hook: result formatting, statistical analysis]
# 5. Output result
Six Hook Points
LangChain's Middleware provides 6 hooks, divided into two categories by execution timing:
| Hook | Execution Frequency | Execution Position | Main Purpose |
|---|---|---|---|
| before_agent | Once | Before Agent starts | Initialization, permission checks, input preprocessing |
| before_model | Every loop iteration | Before model call | Message preprocessing, dynamic context injection |
| wrap_model_call | Every loop iteration | Wraps model call | Retry, degradation, caching, request rewriting |
| after_model | Every loop iteration | After model call | Content moderation, response filtering, logging |
| wrap_tool_call | Every tool call | Wraps tool execution | Tool retry, result caching, parameter rewriting |
| after_agent | Once | After Agent ends | Output formatting, statistics, resource cleanup |
Two Ways to Use
Middleware can be used in two ways: class inheritance or decorators.
Method 1: Decorator (Recommended)
Example
# Decorator approach: simple and intuitive
@before_model
def log_before(state, runtime):
"""Log before each model call"""
msg_count = len(state.get("messages", []))
print(f"[before_model] Current message count: {msg_count}")
return None
@after_model
def log_after(state, runtime):
"""Log after each model call"""
last_msg = state["messages"][-1] if state.get("messages") else None
if last_msg and hasattr(last_msg, 'tool_calls') and last_msg.tool_calls:
print(f"[after_model] Model requested a tool call")
return None
Method 2: Class Inheritance (Suitable for Complex Logic)
Example
class LoggingMiddleware(AgentMiddleware):
"""Custom logging middleware"""
@property
def name(self) -> str:
# Custom middleware name (defaults to class name)
return "logging"
def before_agent(self, state, runtime):
"""Logic before Agent starts"""
print("[Logging] Agent started executing")
return None
def before_model(self, state, runtime):
"""Logic before model call"""
msg_count = len(state.get("messages", []))
print(f"[Logging] Preparing to call model, currently {msg_count} messages")
return None
def after_model(self, state, runtime):
"""Logic after model call"""
print("[Logging] Model call completed")
return None
def after_agent(self, state, runtime):
"""Logic after Agent ends"""
print("[Logging] Agent execution finished")
return None
Complete Lifecycle Example
Example
load_dotenv()
from langchain.agents import create_agent
from langchain.agents.middleware import (
before_agent, after_agent,
before_model, after_model,
)
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
from langchain.tools import tool
@before_agent
def start_log(state, runtime):
"""Before Agent starts"""
print(">>> [before_agent] Agent started <<<")
runtime.stream_writer({"type": "lifecycle", "phase": "start"})
return None
@before_model
def pre_model(state, runtime):
"""Before each model call"""
msg_count = len(state.get("messages", []))
print(f" -> [before_model] Message {msg_count}")
return None
@after_model
def post_model(state, runtime):
"""After each model call"""
last = state["messages"][-1] if state.get("messages") else None
if hasattr(last, 'tool_calls') and last.tool_calls:
tools = [tc['name'] for tc in last.tool_calls]
print(f" <- [after_model] Requested tool: {tools}")
else:
content = str(last.content)[:50] if last and hasattr(last, 'content') else ""
print(f" <- [after_model] Direct reply: {content}...")
return None
@after_agent
def end_log(state, runtime):
"""After Agent ends"""
total_msgs = len(state.get("messages", []))
print(f"<<< [after_agent] Agent ended, {total_msgs} messages in total <<<")
return None
@tool
def get_weather(city: str) -> str:
"""Query weather"""
return f"{city}: Sunny, 25°C"
model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
model=model,
tools=[get_weather],
middleware=[start_log, pre_model, post_model, end_log],
system_prompt="You are an assistant.",
)
print("\n========== First question (requires tool) ==========")
result = agent.invoke({
"messages": [HumanMessage(content="What's the weather in Hangzhou?")]
})
print(f"\nFinal reply: {result['messages'][-1].content}")
print("\n========== Second question (no tool needed) ==========")
result = agent.invoke({
"messages": [HumanMessage(content="Hello")]
})
print(f"\nFinal reply: {result['messages'][-1].content}")
Output:
========== 第一个问题(需要工具) ========== >>> [before_agent] Agent 开始 <<< -> [before_model] 第 2 条消息 <- [after_model] 请求工具: ['get_weather'] -> [before_model] 第 3 条消息 <- [after_model] 直接回复: 杭州今天晴,气温25°C。... <<< [after_agent] Agent 结束,共 4 条消息 <<< Final reply: 杭州今天晴,气温25°C。 ========== 第二个问题(无需工具) ========== >>> [before_agent] Agent 开始 <<< -> [before_model] 第 1 条消息 <- [after_model] 直接回复: 你好!有什么可以帮你的?... <<< [after_agent] Agent 结束,共 2 条消息 <<< Final reply: Hello! How can I help you?
From the output, you can see:
- before_agent and after_agent: executed only once per question
- before_model and after_model: executed on every model call (the first question called the model twice, so each ran twice)
Middleware Return Value
Middleware's return value determines whether to modify the Agent state or control the flow:
| Return Value | Effect | Example |
|---|---|---|
| None | Does not modify any state; continues the normal flow | Logging only |
| dict | Updates Agent state (merged into the current state) | Return {"custom_field": "value"} |
| Dict containing jump_to | Jump to the specified node | Return {"jump_to": "end"} |
Other ExtensionsThe returned dict is merged through the Agent state's reducer. For the messages field, the add_messages reducer is used, so returned messages are appended rather than overwritten.