Pandas df.to_parquet() Function
to_parquet()It is a method of DataFrame, used to export data to a file in Parquet format.
Parquet is a columnar storage file format specifically designed for big data analysis scenarios. It has advantages such as high compression ratio, high read/write performance, and support for complex data types. It is the standard format for big data frameworks such as Apache Hadoop and Apache Spark.
Basic Syntax and Parameters
Syntax Format
DataFrame.to_parquet(path, engine='auto', compression='snappy',
index=None, partition_cols=None, storage_options=None, ...)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| path | str, path object | File path | Required |
| engine | str | Engine: 'auto', 'pyarrow', 'fastparquet' | 'auto' |
| compression | str | Compression: 'snappy', 'gzip', 'brotli', None | 'snappy' |
| index | bool, None | Whether to include the index | None |
| partition_cols | list | Partition column(s), store data partitioned by column | None |
Return Value Description
- Return Type:
None - Writes data directly to a Parquet file, no return value.
Examples
Through the following examples, comprehensively masterto_parquet()the various usages.
Example 1: Basic Usage - Export to a Parquet File
First create a DataFrame, then useto_parquet()to export to a Parquet file.
Example
# Create a sample DataFrame
data = {
'name': ['Tom', 'Jerry', 'Mike', 'Lucy', 'John'],
'age': [28, 35, 42, 26, 31],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen', 'Hangzhou'],
'salary': [8000, 12000, 15000, 7000, 9000],
'department': ['IT', 'HR', 'Sales', 'IT', 'HR']
}
df = pd.DataFrame(data)
# Example 1a: Basic export
# path: file path (required)
# Uses snappy compression by default
df.to_parquet('employees.parquet')
print("Exported to employees.parquet")
# Check file size
import os
file_size = os.path.getsize('employees.parquet')
print(f"File size: {file_size} bytes")
# Example 1b: Read and verify
# Requires installing pyarrow or fastparquet
df_check = pd.read_parquet('employees.parquet')
print("nVerify by reading:")
print(df_check)
Output:
已导出到 employees.parquet
文件大小: 约 600-800 字节(远小于 CSV)
验证读取:
name age city salary department
0 Tom 28 Beijing 8000 IT
1 Jerry 35 Shanghai 12000 HR
2 Mike 42 Guangzhou 15000 Sales
3 Lucy 26 Shenzhen 7000 IT
4 John 31 Hangzhou 9000 HR
Code Explanation:
to_parquet()Export the DataFrame to Parquet format.- By default, uses
snappycompression, with remarkable compression results. - Parquet files are much smaller than CSV files and also faster to read.
Example 2: Selecting Engine and Compression Method
The Parquet format supports multiple engines and compression methods, which can be selected as needed.
Example
import os
# Create a larger DataFrame to observe the compression effect
import numpy as np
np.random.seed(42)
df_large = pd.DataFrame({
'id': range(10000),
'value': np.random.randn(10000),
'category': np.random.choice(['A', 'B', 'C', 'D'], 10000),
'name': np.random.choice(['Tom', 'Jerry', 'Mike', 'Lucy', 'John'], 10000)
})
# Example 2a: Use different compression methods
# snappy: fast compression, high speed, moderate compression ratio (default)
df_large.to_parquet('output_snappy.parquet', compression='snappy')
# gzip: high compression ratio, smaller files
df_large.to_parquet('output_gzip.parquet', compression='gzip')
# brotli: higher compression ratio
df_large.to_parquet('output_brotli.parquet', compression='brotli')
# no compression
df_large.to_parquet('output_none.parquet', compression=None)
# Compare file sizes
print("Comparison of file sizes with different compression methods:")
for name in ['snappy', 'gzip', 'brotli', 'none']:
size = os.path.getsize(f'output_{name}.parquet')
print(f" {name}: {size:,} bytes")
print()
# Example 2b: Choose the engine
# auto: automatic selection (default)
# pyarrow: Apache Arrow implementation, full-featured, good performance
# fastparquet: pure Python implementation, good compatibility
print("Available engines: auto, pyarrow, fastparquet")
print("Currently using:", end=" ")
# Check available engines
try:
import pyarrow
print("pyarrow")
except ImportError:
pass
try:
import fastparquet
print("fastparquet")
except ImportError:
pass
Output:
不同压缩方式的文件大小对比: snappy: 约 100KB gzip: 约 80KB brotli: 约 70KB none: 约 200KB 不同压缩方式各有优劣: snappy: 速度快,压缩比适中(默认推荐) gzip: 压缩比更高,适合存储 brotli: 最高压缩比,适合冷数据 none: 无压缩,速度最快
Code Explanation:
compressionThe compression parameter allows selecting different compression methods.snappyis the default option, balancing speed and compression ratio.gzipHas a higher compression ratio, but is slightly slower.engineThe engine parameter allows selecting which engine to use.
Example 3: Partitioned Storage
Parquet supports partitioned storage by column, which is an important feature in big data analysis.
Example
import os
import shutil
# Create a DataFrame
df = pd.DataFrame({
'name': ['Tom', 'Jerry', 'Mike', 'Lucy', 'John', 'Mary', 'Bob', 'Alice'],
'age': [28, 35, 42, 26, 31, 29, 38, 24],
'department': ['IT', 'HR', 'Sales', 'IT', 'HR', 'IT', 'Sales', 'HR'],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen', 'Hangzhou',
'Beijing', 'Shanghai', 'Beijing']
})
# Example 3a: Partition by a single column
# partition_cols specifies the partition column, which generates subdirectories
if os.path.exists('partitioned'):
shutil.rmtree('partitioned')
df.to_parquet('partitioned', partition_cols=['department'])
print("Exported with partitioning by department")
# View the partition directory structure
for root, dirs, files in os.walk('partitioned'):
level = root.replace('partitioned', '').count(os.sep)
indent = ' ' * 2 * level
print(f'{indent}{os.path.basename(root)}/')
subindent = ' ' * 2 * (level + 1)
for file in files:
print(f'{subindent}{file}')
Output:
已按 department 分区导出
分区目录结构:
partitioned/
department=HR/
xxx.parquet
department=IT/
xxx.parquet
department=Sales/
xxx.parquet
Code Explanation:
partition_colsThe partition_cols parameter performs partitioned storage by the specified column.- Partitioned storage creates subdirectories in the file system, one directory per partition value.
- Partitioned storage is very beneficial for big data queries, as only the needed partitions need to be read.
Example 4: Handling the Index
When exporting to Parquet, you can choose whether to include the index.
Example
# Create a DataFrame with an index
df = pd.DataFrame({
'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
'age': [28, 35, 42, 26],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen']
})
df.index = ['A001', 'A002', 'A003', 'A004']
# Example 4a: Include index by default
df.to_parquet('with_index.parquet')
print("Export (index included by default)")
# Example 4b: Exclude index
df.to_parquet('without_index.parquet', index=False)
print("Export (index not included)")
# Example 4c: Explicitly include index
df.to_parquet('explicit_index.parquet', index=True)
print("Export (explicitly include index)")
# Read and compare
print("nRead and compare the differences:")
print("nOriginal data:")
print(df)
print("nRead with_index.parquet:")
print(pd.read_parquet('with_index.parquet'))
print("nRead without_index.parquet:")
print(pd.read_parquet('without_index.parquet'))
print("nRead explicit_index.parquet:")
print(pd.read_parquet('explicit_index.parquet'))
Output:
导出(默认包含索引)
导出(不包含索引)
导出(显式包含索引)
读取并比较差异:
原始数据:
name age city
A001 Tom 28 Beijing
A002 Jerry 35 Shanghai
A003 Mike 42 Guangzhou
A004 Lucy 26 Shenzhen
读取 with_index.parquet:
name age city index
0 Tom 28 Beijing A001
1 Jerry 35 Shanghai A002
2 Mike 42 Guangzhou A003
3 Lucy 26 Shenzhen A004
读取 without_index.parquet:
name age city
0 Tom 28 Beijing
1 Jerry 35 Shanghai
2 Mike 42 Guangzhou
3 Lucy 26 Shenzhen
读取 explicit_index.parquet:
name age city index
0 Tom 28 Beijing A001
1 Jerry 35 Shanghai A002
2 Mike 42 Guangzhou A003
3 Lucy 26 Shenzhen A004
Code Explanation:
index=TrueExplicitly includes the index as a column.index=FalseDoes not include the index.- The default behavior depends on pandas' index settings.
Notes
- Using
to_parquet()requires installingpyarroworfastparquet。 - Recommended to install
pyarrow:pip install pyarrow。 - Parquet is columnar storage, suitable for big data analysis scenarios, but not suitable for small data.
- Partitioned storage can significantly improve big data query performance.
- Parquet supports complex data types (nested structures), but they are not commonly used in pandas DataFrames.
Summary
to_parquet()It is a method for DataFrame to export to Parquet format. Parquet is the standard columnar storage format in the big data field, with advantages such as high compression ratio, high performance, and support for partitioning.
In big data processing scenarios, Parquet is the preferred data format. It is perfectly compatible with big data frameworks such as Apache Spark and Apache Hive. When processing large-scale data, it is recommended to use Parquet format instead of CSV or Excel.
Other Extensions
Pandas Common Functions