LangChain Structured Output

Most of the time, what you need is not a piece of free text, but structured data—such as a JSON object.

LangChain Structured Output lets the Agent return results in the format you specify, making it easy for programs to use directly.


Why Do You Need Structured Output?

Suppose you need to extract a name, age, and occupation from a user description:

MethodOutput FormatSubsequent Processing
Normal ReplyZhang San is 28 years old and is an engineer.Requires regex or calling the model again to parse
Structured Output{name: "Zhang San", age: 28, job: "Engineer"}Used directly as a Python object

Structured output saves the step of "parsing data from text", allowing AI output to be used directly by programs.


The Simplest Usage — Passing in a Pydantic Model

Just pass the Pydantic model to the response_format parameter:

Example

from dotenv import load_dotenv
load_dotenv()

from pydantic import BaseModel, Field
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage


# Define the expected output structure
class CourseInfo(BaseModel):
    """Example Tutorial course extraction result"""
    course_name: str = Field(description=Course name)
    difficulty: str = Field(description=Difficulty: Beginner/Intermediate/Advanced)
    estimated_hours: int = Field(description=Estimated study duration (hours))
    is_free: bool = Field(description=Whether it is free)


model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
    model=model,
    response_format=CourseInfo,  # Pass in the Pydantic model
    system_prompt=You are the course assistant at Example Tutorial. Extract course information from user descriptions.,
)

# The user enters an unstructured description
result = agent.invoke({
    "messages": [HumanMessage(
        content=I've recently been studying the Python3 Basics Tutorial, it's beginner-level,
                It takes about 20 hours to study, and it's completely free.
    )]
})

# Get the structured result from structured_response
if "structured_response" in result:
    course = result["structured_response"]
    print(fCourse name: {course.course_name})
    print(fDifficulty: {course.difficulty})
    print(fEstimated duration: {course.estimated_hours} hours)
    print(fFree: {'Yes' if course.is_free else 'No'})
    print(fObject type: {type(course)})

Run result:

Course name: Python3 基础教程
Difficulty: 入门
Duration: 20 小时
Free: 是
Object type: <class '__main__.CourseInfo'>

The returned structured_response is a Pydantic model instance, not a plain dictionary. This means you can access attributes such as .course_name, and the IDE can provide auto-completion.


Structured Output Coexisting with Tools

response_format and tools can be used together—the Agent calls tools when needed and finally outputs structured data:

Example

from pydantic import BaseModel, Field
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
from langchain.tools import tool


@tool
def search_course(keyword: str) -> str:
    """Search course information on Example Tutorial"""
    courses = {
        "python": Python3 Basics Tutorial | Beginner | Free | 30 chapters | About 20 hours,
        "java": Java Basics Tutorial | Beginner | Free | 35 chapters | About 25 hours,
        Data Analysis: Python Data Analysis | Intermediate | Member | 25 chapters | About 30 hours,
    }
    return courses.get(keyword.lower(), fNo course found for '{keyword}')


class CourseRecommendation(BaseModel):
    """Course recommendation result"""
    course_name: str = Field(description=Recommended course name)
    reason: str = Field(description=Recommendation reason)
    difficulty: str = Field(description=Difficulty: Beginner/Intermediate/Advanced)


model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
    model=model,
    tools=[search_course],
    response_format=CourseRecommendation,
    system_prompt=You are the course consultant at Example Tutorial. Search for courses first, then give recommendations.,
)

result = agent.invoke({
    "messages": [HumanMessage(content=I want to learn Python. What do you recommend?)]
})

rec = result["structured_response"]
print(fRecommended course: {rec.course_name})
print(fRecommendation reason: {rec.reason})
print(fDifficulty: {rec.difficulty})

# View the full process
print("\n"=== Execution Process ===")
for msg in result["messages"]:
    if msg.type == "tool":
        print(f" Call {msg.name}: {msg.content}")

Run result:

Recommended course: Python3 基础教程
Reason: 该课程免费且适合Python初学者,学习时长约20小时
Difficulty: 入门

=== 执行过程 ===
  调用 search_course: Python3 基础教程 | 入门 | 免费 | 30章 | 约20小时

Complex Nested Structures

Pydantic supports complex structures such as nesting, lists, enums, etc.:

Example

from pydantic import BaseModel, Field
from typing import Literal
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage


class Topic(BaseModel):
    """Topic"""
    name: str = Field(description=Topic name)
    order: int = Field(description=Learning order, starting from 1)
    minutes: int = Field(description=Recommended study minutes)


class LearningPlan(BaseModel):
    """Study Plan"""
    goal: str = Field(description=Study goal overview)
    level: Literal[Beginner, Intermediate, Advanced] = Field(description=Difficulty level)
    total_hours: float = Field(description=Total duration (hours))
    topics: list[Topic] = Field(description=Topic list)


model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
    model=model,
    response_format=LearningPlan,
    system_prompt=You are the study planner at Example Tutorial.,
)

result = agent.invoke({
    "messages": [HumanMessage(
        content=Help me create a Python beginner study plan, keeping the total time within 10 hours
    )]
})

plan = result["structured_response"]
print(fGoal: {plan.goal})
print(fDifficulty: {plan.level})
print(fTotal duration: {plan.total_hours} hours)
print(f"\n"Topic list ({len(plan.topics)} items):")
for topic in plan.topics:
    print(f" {topic.order}. {topic.name} ({topic.minutes} minutes)")

Run result:

目标: 掌握 Python 基础语法,能够独立编写简单的 Python 程序
Difficulty: 入门
总时长: 9.5 小时

知识点列表 (6 个):
  1. 环境搭建与基础语法 (60分钟)
  2. 数据类型与变量 (90分钟)
  3. 条件判断与循环 (120分钟)
  4. 函数与模块 (120分钟)
  5. 列表与字典 (90分钟)
  6. 综合练习 (90分钟)

Getting Structured Output from Messages

If you don't need the Agent's tool-calling ability and just want to extract structured information from text, you can use the model directly:

Example

from pydantic import BaseModel, Field
from langchain.chat_models import init_chat_model


class SentimentResult(BaseModel):
    """Sentiment Analysis Result"""
    sentiment: str = Field(description=Positive/Negative/Neutral)
    score: float = Field(description=Sentiment intensity 0~1)
    keywords: list[str] = Field(description=Key sentiment words)


model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)

# Use with_structured_output() directly on the model
# No need for an Agent
structured_model = model.with_structured_output(SentimentResult)

texts = [
    Example Tutorial is really great to use. Highly recommended!,
    This tutorial has too little content. It's not worth it.,
    The weather is nice today.,
]

for text in texts:
    result = structured_model.invoke(text)
    print(fText: {text[:30]}...)
    print(f" Sentiment: {result.sentiment}, Intensity: {result.score}, Keywords: {result.keywords}")

Run result:

Text: Example 真的太好用了,强烈推荐!...
  Sentiment: 积极, 强度: 0.95, 关键词: ['好用', '推荐']
Text: 这个教程内容太少了,不太值。...
  Sentiment: 消极, 强度: 0.7, 关键词: ['内容太少', '不太值']
Text: 今天天气不错。...
  Sentiment: 中性, 强度: 0.1, 关键词: ['不错']

with_structured_output() is a method of the model and can be used without an Agent. If your scenario is 'information extraction' rather than 'multi-step reasoning', using with_structured_output() directly is simpler and more efficient.

Other Extensions