Pandas df.to_json() Function
to_json()It is a method of DataFrame, used to export data to JSON (JavaScript Object Notation) format.
JSON is a standard format for data exchange in Web applications and APIs,to_json()It supports output in multiple JSON formats, including record format, index format, column format, etc. It can convert a pandas DataFrame to a JSON string or write it to a file, facilitating data exchange with other systems.
Basic Syntax and Parameters
Syntax Format
DataFrame.to_json(path_or_buf=None, orient=None, date_format=None,
double_precision=10, force_ascii=True, date_unit='ms',
default_handler=None, lines=False, index=False,
indent=None, ...)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| path_or_buf | str, path object, file-like object | File path; returns a string if None | None |
| orient | str | JSON format: 'split', 'records', 'index', 'columns', 'values', 'table' | None |
| date_format | str | Date format: 'epoch', 'iso' | None |
| double_precision | int | Floating point precision | 10 |
| force_ascii | bool | Whether to force ASCII encoding | True |
| date_unit | str | Date unit: 's', 'ms', 'us', 'ns' | 'ms' |
| lines | bool | JSON Lines format, one JSON object per line | False |
| index | bool | Whether to include index | False |
Return Value
- Return Type:
Noneorstr - When a file path is specified, it writes to the file and returns None.
- When
path_or_buf=NoneWhen not specified, returns a JSON format string.
Examples
Through the following examples, comprehensively masterto_json()the various usages.
Example 1: Basic Usage - Export to JSON String
First, create a DataFrame, then useto_json()to export it to JSON format.
Example
# Create a sample DataFrame
data = {
'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
'age': [28, 35, 42, 26],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen'],
'salary': [8000, 12000, 15000, 7000]
}
df = pd.DataFrame(data)
# Example 1a: The most basic export (no path specified, returns a string)
# Default orient='columns'
json_string = df.to_json()
print("Default JSON format (columns):")
print(json_string)
print()
# Example 1b: Export to file
df.to_json('output_basic.json', orient='records', indent=2)
print("Exported to output_basic.json")
# Read and verify
df_check = pd.read_json('output_basic.json', orient='records')
print("nVerify reading:")
print(df_check)
Expected output:
默认 JSON 格式 (columns):
{"name":{"0":"Tom","1":"Jerry","2":"Mike","3":"Lucy"},"age":{"0":28,"1":35,"2":42,"3":26},...}
默认 JSON 格式 (records):
[
{"name":"Tom","age":28,"city":"Beijing","salary":8000},
{"name":"Jerry","age":35,"city":"Shanghai","salary":12000},
{"name":"Mike","age":42,"city":"Guangzhou","salary":15000},
{"name":"Lucy","age":26,"city":"Shenzhen","salary":7000}
]
Code explanation:
to_json()When no path is specified, a JSON string is returned.- By default
orient='columns', one JSON object per column. indent=2Makes the output easier to read (formatted).
Example 2: Different JSON Formats
JSON data has multiple formats,to_json()and supports multiple output formats to meet different needs.
Example
# Create DataFrame
df = pd.DataFrame({
'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
'age': [28, 35, 42, 26],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen']
})
# Example 2a: records format - array format, one object per row
# Most commonly used, suitable for most APIs
json_records = df.to_json(orient='records')
print("Records format (one object per row):")
print(json_records)
print()
# Example 2b: index format - uses index as key
json_index = df.to_json(orient='index')
print("Index format (uses index as key):")
print(json_index)
print()
# Example 2c: columns format - uses column names as keys (default)
json_columns = df.to_json(orient='columns')
print("Columns format (uses column names as keys):")
print(json_columns)
print()
# Example 2d: split format - split structure
json_split = df.to_json(orient='split')
print("Split format (split structure):")
print(json_split)
print()
# Example 2e: values format - only contains value arrays
json_values = df.to_json(orient='values')
print("Values format (values only):")
print(json_values)
Expected output:
Records 格式(每行一个对象):
[{"name":"Tom","age":28,"city":"Beijing"},
{"name":"Jerry","age":35,"city":"Shanghai"},
{"name":"Mike","age":42,"city":"Guangzhou"},
{"name":"Lucy","age":26,"city":"Shenzhen"}]
Index 格式(以索引为键):
{"0":{"name":"Tom","age":28,"city":"Beijing"},
"1":{"name":"Jerry","age":35,"city":"Shanghai"},
...}
Columns 格式(以列名为键):
{"name":{"0":"Tom","1":"Jerry"...},"age":{"0":28...}...}
Split 格式(分离结构):
{"index":[0,1,2,3],"columns":["name","age","city"],
"data":[["Tom",28,"Beijing"],["Jerry",35,"Shanghai"]...]}
Values 格式(只包含值):
[["Tom",28,"Beijing"],["Jerry",35,"Shanghai"],["Mike",42,"Guangzhou"],["Lucy",26,"Shenzhen"]]
Code explanation:
orient='records': Array format, each row is a JSON object, most commonly used.orient='index': Uses the DataFrame index as the top-level key.orient='columns': Uses column names as keys, the default format.orient='split': Split structure, easy for programs to parse.orient='values': Pure numeric array, no column names or index.
Example 3: Handling Dates and Chinese Encoding
When exporting JSON, dates and Chinese encoding need to be handled correctly.
Example
from datetime import datetime
# Example 3a: Handle date fields
df_date = pd.DataFrame({
'name': ['Tom', 'Jerry'],
'birthday': [datetime(1995, 3, 15), datetime(1988, 7, 22)],
'join_date': [datetime(2020, 1, 10), datetime(2019, 3, 5)]
})
# Default date format is timestamp (milliseconds)
json_timestamp = df_date.to_json(date_format='epoch', orient='records')
print("Date format - timestamp:")
print(json_timestamp)
print()
# ISO format date
json_iso = df_date.to_json(date_format='iso', orient='records')
print("Date format - ISO:")
print(json_iso)
print()
# Example 3b: Handle Chinese
df_cn = pd.DataFrame({
'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
'age': [28, 35, 42],
'city': ['Beijing', 'Shanghai', 'Guangzhou']
})
# force_ascii=False allows non-ASCII characters
json_cn = df_cn.to_json(orient='records', force_ascii=False)
print("Chinese content (force_ascii=False):")
print(json_cn)
# Write to file (Chinese supported by default)
df_cn.to_json('output_cn.json', orient='records', force_ascii=False)
print("nExported to output_cn.json")
# Read and verify
df_cn_check = pd.read_json('output_cn.json')
print("Verify reading:")
print(df_cn_check)
Expected output:
日期格式 - 时间戳:
[{"name":"Tom","birthday":796...
{"name":"Jerry","birthday":584...}
日期格式 - ISO:
[{"name":"tom","birthday":"1995-03-15T00:00:00.000Z",...}]
中文内容(force_ascii=False):
[{"姓名":"张三","年龄":28,"城市":"北京"},
{"姓名":"李四","年龄":35,"城市":"上海"},
{"姓名":"王五","年龄":42,"城市":"广州"}]
已导出到 output_cn.json
验证读取:
姓名 年龄 城市
0 张三 (read_json 的 格式 orient='records' 匹配 to_json 的 orient='records') 北京
1 配对 李四 35 上海
2 王五 42 广州
</对>
Code explanation:
date_format='epoch'Converts dates to timestamp (milliseconds).date_format='iso'Converts dates to ISO 8601 format.force_ascii=FalseAllows Chinese characters to be preserved in JSON instead of being escaped to Unicode.
Example 4: JSON Lines Format
JSON Lines is a format with one JSON object per line, commonly used for logging and big data processing.
Example
# Create DataFrame
df = pd.DataFrame({
'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
'age': [28, 35, 42, 26],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen']
})
# Example 4a: JSON Lines format (one JSON object per line)
# lines=True enables JSON Lines format
json_lines = df.to_json(orient='records', lines=True)
print("JSON Lines format:")
print(json_lines)
print()
# Write to file
df.to_json('output_lines.json', orient='records', lines=True)
print("Exported to output_lines.json")
# Example 4b: Read JSON Lines format
# Need to read line by line
import json
df_lines_list = []
with open('output_lines.json', 'r', encoding='utf-8') as f:
for line in f:
df_lines_list.append(json.loads(line.strip()))
df_from_lines = pd.DataFrame(df_lines_list)
print("nRead JSON Lines:")
print(df_from_lines)
Expected output:
JSON Lines 格式:
{"name":"Tom","age":28,"city":"Beijing"}
{"name":"Jerry","age":35,"city":"Shanghai"}
{"name":"Mike","age":42,"city":"city":"Guangzhou"}
{"name":"Lucy","age":26,"city":"Shenzhen"}
已导出到 output_lines.json
读取 JSON Lines:
name age city
0 Tom 28 Beijing
对 读 Jerry 35 Shanghai
2 Mike 42 格式 Guangzhou
3 Lucy 26 Shenzhen
Code explanation:
lines=TrueEnables JSON Lines format, one JSON object per line.- JSON Lines is suitable for streaming processing and big data scenarios.
- When reading, it needs to be parsed line by line and converted to a DataFrame.
Notes
- By default
orient='columns', it is generally recommended to useorient='records'more universal. - Use
force_ascii=Falsecan preserve Chinese characters. - JSON Lines format needs to be used with
orient='records'for use. read_json()When reading, you need to use the sameorientparameter.- It is recommended to use date format
date_format='iso', which is more readable.
Summary
to_json()It is the core method for DataFrame to export to JSON format. It supports multiple JSON formats and can meet the needs of different scenarios such as Web API and data exchange.
In actual work, JSON is the most commonly used data format in Web applications,to_json()combined withread_json()can realize data import and export.orientthe different formats of parameters, and the pairing withread_json()for use.
Pandas Common Functions