Pandas df.to_json() Function

Pandas 常用函数Pandas Common Functions


to_json()It is a method of DataFrame, used to export data to JSON (JavaScript Object Notation) format.

JSON is a standard format for data exchange in Web applications and APIs,to_json()It supports output in multiple JSON formats, including record format, index format, column format, etc. It can convert a pandas DataFrame to a JSON string or write it to a file, facilitating data exchange with other systems.


Basic Syntax and Parameters

Syntax Format

DataFrame.to_json(path_or_buf=None, orient=None, date_format=None,
                  double_precision=10, force_ascii=True, date_unit='ms',
                  default_handler=None, lines=False, index=False,
                  indent=None, ...)

Parameter Description

ParameterTypeDescriptionDefault Value
path_or_bufstr, path object, file-like objectFile path; returns a string if NoneNone
orientstrJSON format: 'split', 'records', 'index', 'columns', 'values', 'table'None
date_formatstrDate format: 'epoch', 'iso'None
double_precisionintFloating point precision10
force_asciiboolWhether to force ASCII encodingTrue
date_unitstrDate unit: 's', 'ms', 'us', 'ns''ms'
linesboolJSON Lines format, one JSON object per lineFalse
indexboolWhether to include indexFalse

Return Value

  • Return Type:Noneorstr
  • When a file path is specified, it writes to the file and returns None.
  • Whenpath_or_buf=NoneWhen not specified, returns a JSON format string.

Examples

Through the following examples, comprehensively masterto_json()the various usages.

Example 1: Basic Usage - Export to JSON String

First, create a DataFrame, then useto_json()to export it to JSON format.

Example

import pandas as pd

# Create a sample DataFrame
data = {
    'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
    'age': [28, 35, 42, 26],
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen'],
    'salary': [8000, 12000, 15000, 7000]
}
df = pd.DataFrame(data)

# Example 1a: The most basic export (no path specified, returns a string)
# Default orient='columns'
json_string = df.to_json()
print("Default JSON format (columns):")
print(json_string)
print()

# Example 1b: Export to file
df.to_json('output_basic.json', orient='records', indent=2)
print("Exported to output_basic.json")

# Read and verify
df_check = pd.read_json('output_basic.json', orient='records')
print("nVerify reading:")
print(df_check)

Expected output:

默认 JSON 格式 (columns):
{"name":{"0":"Tom","1":"Jerry","2":"Mike","3":"Lucy"},"age":{"0":28,"1":35,"2":42,"3":26},...}

默认 JSON 格式 (records):
[
  {"name":"Tom","age":28,"city":"Beijing","salary":8000},
  {"name":"Jerry","age":35,"city":"Shanghai","salary":12000},
  {"name":"Mike","age":42,"city":"Guangzhou","salary":15000},
  {"name":"Lucy","age":26,"city":"Shenzhen","salary":7000}
]

Code explanation:

  • to_json()When no path is specified, a JSON string is returned.
  • By defaultorient='columns', one JSON object per column.
  • indent=2Makes the output easier to read (formatted).

Example 2: Different JSON Formats

JSON data has multiple formats,to_json()and supports multiple output formats to meet different needs.

Example

import pandas as pd

# Create DataFrame
df = pd.DataFrame({
    'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
    'age': [28, 35, 42, 26],
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen']
})

# Example 2a: records format - array format, one object per row
# Most commonly used, suitable for most APIs
json_records = df.to_json(orient='records')
print("Records format (one object per row):")
print(json_records)
print()

# Example 2b: index format - uses index as key
json_index = df.to_json(orient='index')
print("Index format (uses index as key):")
print(json_index)
print()

# Example 2c: columns format - uses column names as keys (default)
json_columns = df.to_json(orient='columns')
print("Columns format (uses column names as keys):")
print(json_columns)
print()

# Example 2d: split format - split structure
json_split = df.to_json(orient='split')
print("Split format (split structure):")
print(json_split)
print()

# Example 2e: values format - only contains value arrays
json_values = df.to_json(orient='values')
print("Values format (values only):")
print(json_values)

Expected output:

Records 格式(每行一个对象):
[{"name":"Tom","age":28,"city":"Beijing"},
{"name":"Jerry","age":35,"city":"Shanghai"},
{"name":"Mike","age":42,"city":"Guangzhou"},
{"name":"Lucy","age":26,"city":"Shenzhen"}]

Index 格式(以索引为键):
{"0":{"name":"Tom","age":28,"city":"Beijing"},
"1":{"name":"Jerry","age":35,"city":"Shanghai"},
...}

Columns 格式(以列名为键):
{"name":{"0":"Tom","1":"Jerry"...},"age":{"0":28...}...}

Split 格式(分离结构):
{"index":[0,1,2,3],"columns":["name","age","city"],
"data":[["Tom",28,"Beijing"],["Jerry",35,"Shanghai"]...]}

Values 格式(只包含值):
[["Tom",28,"Beijing"],["Jerry",35,"Shanghai"],["Mike",42,"Guangzhou"],["Lucy",26,"Shenzhen"]]

Code explanation:

  • orient='records': Array format, each row is a JSON object, most commonly used.
  • orient='index': Uses the DataFrame index as the top-level key.
  • orient='columns': Uses column names as keys, the default format.
  • orient='split': Split structure, easy for programs to parse.
  • orient='values': Pure numeric array, no column names or index.

Example 3: Handling Dates and Chinese Encoding

When exporting JSON, dates and Chinese encoding need to be handled correctly.

Example

import pandas as pd
from datetime import datetime

# Example 3a: Handle date fields
df_date = pd.DataFrame({
    'name': ['Tom', 'Jerry'],
    'birthday': [datetime(1995, 3, 15), datetime(1988, 7, 22)],
    'join_date': [datetime(2020, 1, 10), datetime(2019, 3, 5)]
})

# Default date format is timestamp (milliseconds)
json_timestamp = df_date.to_json(date_format='epoch', orient='records')
print("Date format - timestamp:")
print(json_timestamp)
print()

# ISO format date
json_iso = df_date.to_json(date_format='iso', orient='records')
print("Date format - ISO:")
print(json_iso)
print()

# Example 3b: Handle Chinese
df_cn = pd.DataFrame({
    'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
    'age': [28, 35, 42],
    'city': ['Beijing', 'Shanghai', 'Guangzhou']
})

# force_ascii=False allows non-ASCII characters
json_cn = df_cn.to_json(orient='records', force_ascii=False)
print("Chinese content (force_ascii=False):")
print(json_cn)

# Write to file (Chinese supported by default)
df_cn.to_json('output_cn.json', orient='records', force_ascii=False)
print("nExported to output_cn.json")

# Read and verify
df_cn_check = pd.read_json('output_cn.json')
print("Verify reading:")
print(df_cn_check)

Expected output:

日期格式 - 时间戳:
[{"name":"Tom","birthday":796...
{"name":"Jerry","birthday":584...}

日期格式 - ISO:
[{"name":"tom","birthday":"1995-03-15T00:00:00.000Z",...}]

中文内容(force_ascii=False):
[{"姓名":"张三","年龄":28,"城市":"北京"},
{"姓名":"李四","年龄":35,"城市":"上海"},
{"姓名":"王五","年龄":42,"城市":"广州"}]

已导出到 output_cn.json

验证读取:
    姓名   年龄   城市
0  张三   (read_json 的    格式    orient='records' 匹配 to_json 的 orient='records')    北京
1 配对    李四   35    上海
2  王五   42    广州
</对>

Code explanation:

  • date_format='epoch'Converts dates to timestamp (milliseconds).
  • date_format='iso'Converts dates to ISO 8601 format.
  • force_ascii=FalseAllows Chinese characters to be preserved in JSON instead of being escaped to Unicode.

Example 4: JSON Lines Format

JSON Lines is a format with one JSON object per line, commonly used for logging and big data processing.

Example

import pandas as pd

# Create DataFrame
df = pd.DataFrame({
    'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
    'age': [28, 35, 42, 26],
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen']
})

# Example 4a: JSON Lines format (one JSON object per line)
# lines=True enables JSON Lines format
json_lines = df.to_json(orient='records', lines=True)
print("JSON Lines format:")
print(json_lines)
print()

# Write to file
df.to_json('output_lines.json', orient='records', lines=True)
print("Exported to output_lines.json")

# Example 4b: Read JSON Lines format
# Need to read line by line
import json
df_lines_list = []
with open('output_lines.json', 'r', encoding='utf-8') as f:
    for line in f:
        df_lines_list.append(json.loads(line.strip()))

df_from_lines = pd.DataFrame(df_lines_list)
print("nRead JSON Lines:")
print(df_from_lines)

Expected output:

JSON Lines 格式:
{"name":"Tom","age":28,"city":"Beijing"}
{"name":"Jerry","age":35,"city":"Shanghai"}
{"name":"Mike","age":42,"city":"city":"Guangzhou"}
{"name":"Lucy","age":26,"city":"Shenzhen"}

已导出到 output_lines.json

读取 JSON Lines:
    name  age       city
0    Tom   28    Beijing
对    读    Jerry   35    Shanghai
2   Mike   42   格式    Guangzhou
3   Lucy   26    Shenzhen

Code explanation:

  • lines=TrueEnables JSON Lines format, one JSON object per line.
  • JSON Lines is suitable for streaming processing and big data scenarios.
  • When reading, it needs to be parsed line by line and converted to a DataFrame.

Notes

  • By defaultorient='columns', it is generally recommended to useorient='records'more universal.
  • Useforce_ascii=Falsecan preserve Chinese characters.
  • JSON Lines format needs to be used withorient='records'for use.
  • read_json()When reading, you need to use the sameorientparameter.
  • It is recommended to use date formatdate_format='iso', which is more readable.

Summary

to_json()It is the core method for DataFrame to export to JSON format. It supports multiple JSON formats and can meet the needs of different scenarios such as Web API and data exchange.

In actual work, JSON is the most commonly used data format in Web applications,to_json()combined withread_json()can realize data import and export.orientthe different formats of parameters, and the pairing withread_json()for use.

Pandas 常用函数Pandas Common Functions

Other Extensions