Pandas df.fillna() Function
df.fillna()is a function in Pandas used to fill missing values.
anddropna()Unlike deleting missing values,fillna()it allows us to fill missing data with specified values, mean, median, forward fill, or backward fill, etc. This is very useful for maintaining data integrity and analysis accuracy.
Basic Syntax and Parameters
fillna()is a member function of DataFrame, called via the dot operator.to call.
Syntax Format
DataFrame.fillna(value=None, method=None, axis=None, inplace=False, limit=None, downcast=None)
Parameter Description
| Parameter | Type | Required | Description | Default Value |
|---|---|---|---|---|
| value | scalar, dict, Series, DataFrame | Optional | The value used to fill missing values. It can be a constant, a dictionary (specifying different values for different columns), a Series, or a DataFrame. | None |
| method | str | Optional | Fill method.'ffill'or'pad'Indicates filling with the previous value (forward fill);'bfill'or'backfill'Indicates filling with the next value (backward fill). |
None |
| axis | int or str | Optional | Specifies the fill direction.0or'index'Indicates filling along rows;1or'columns'Indicates filling along columns. Only whenmethodis notNoneis it valid. |
0 |
| inplace | bool | Optional | IfTrue, it directly modifies the original DataFrame and does not return a new object; ifFalse, it returns a new DataFrame, and the original data remains unchanged. |
False |
| limit | int | Optional | Specifies the maximum number of consecutive fills. For example, settinglimit=1means filling only 1 missing value at a time. |
None |
| downcast | dict or str | Optional | Specifies the rules for downcasting data types, such as converting float64 to int64. | None |
Return Value Description
- Returns a new DataFrame (if
inplace=False), orNone(ifinplace=True)。 - In the returned DataFrame, missing values have been filled.
Examples
Let's, through a series of examples, thoroughly masterfillna()the usage of df.fillna().
Example 1: Fill All Missing Values with a Constant
The simplest usage is to fill all missing values with a constant.
Example
import numpy as np
# Create a DataFrame containing missing values
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'Age': [25, np.nan, 35, np.nan], # missing value
'Salary': [5000, 6000, np.nan, 8000], # missing value
'Department': ['Technology', 'Marketing', 'Technology', np.nan] # missing value
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Fill all missing values with the constant 0
df_filled = df.fillna(0)
print("Data after filling with 0:")
print(df_filled)
Expected output:
原始数据:
姓名 年龄 薪资 部门
0 张三 25.0 5000.0 技术
1 李四 NaN 6000.0 市场
2 王五 35.0 NaN 技术
3 赵六 NaN 8000.0 NaN
==================================================
用0填充后的数据:
名称 年龄 薪资 部门
0 张三 25.0 5000.0 技术
1 李四 0.0 6000.0 市场
2 王五 35.0 0.0 技术
3 赵六 0.0 8000.0 0
Code explanation:
- We created a DataFrame containing multiple missing values.
- Use
df.fillna(0)to replace all missing values with 0. - This method is simple and direct, but may not be suitable for all scenarios; for example, filling age and salary with 0 may not be reasonable.
Example 2: Specify Different Fill Values for Different Columns with a Dictionary
You can specify different fill values for different columns to make the data more reasonable.
Example
import numpy as np
# Create a DataFrame containing missing values
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'Age': [25, np.nan, 35, np.nan],
'Salary': [5000, 6000, np.nan, 8000],
'Department': ['Technology', 'Marketing', 'Technology', np.nan],
'Performance': [85, 90, np.nan, 95]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Use a dictionary to specify different fill values for different columns
fill_values = {
'Age': 30, # Fill missing ages with 30
'Salary': 6500, # Fill missing salaries with 6500
'Department': 'Unassigned', # Fill missing departments with "Unassigned"
'Performance': 80 # Fill missing performance with 80
}
df_filled = df.fillna(fill_values)
print("Data after filling different columns with different values:")
print(df_filled)
Expected output:
原始数据:
姓名 年龄 薪资 部门 绩效
0 张三 25.0 5000.0 技术 85
1 李四 NaN 6000.0 市场 90
2 王五 35.0 NaN 技术 NaN
3 赵六 NaN 8000.0 NaN 95
==================================================
不同列用不同值填充后的数据:
姓名 年龄 薪资 部门 绩效
0 张三 25.0 5000.0 技术 85
1 李四 30.0 6000.0 市场 90
2 王五 35.0 6500.0 技术 80
3 赵六 30.0 8000.0 未分配 95
Code explanation:
- Through the dictionary
fill_values, we specified reasonable fill values for each column. - This method is more flexible and allows setting appropriate default values for different fields based on business logic.
Example 3: Fill Numeric Columns with Mean or Median
For numeric data, filling with the mean or median is a common and reasonable method.
Example
import numpy as np
# Create a DataFrame containing missing values
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
'Math': [85, 90, np.nan, 78, 92],
'English': [88, np.nan, 82, 85, 90],
'Physics': [np.nan, 85, 90, 88, 95]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Calculate the mean of each column
math_mean = df['Math'].mean()
english_mean = df['English'].mean()
physics_mean = df['Physics'].mean()
print(f"Math mean: {math_mean:.2f}")
print(f"English mean: {english_mean:.2f}")
print(f"Physics mean: {physics_mean:.2f}")
print("=" * 50)
# Fill with mean
df_filled = df.fillna({
'Math': math_mean,
'English': english_mean,
'Physics': physics_mean
})
print("Data after filling with mean:")
print(df_filled)
Expected output:
原始数据:
姓名 数学 英语 物理
0 张三 85.0 88.0 NaN
1 李四 90.0 NaN 85.0
2 王五 NaN 82.0 90.0
3 赵六 78.0 85.0 88.0
4 钱七 92.0 90.0 95.0
==================================================
数学均值: 86.25
英语均值: 86.25
物理均值: 89.50
==================================================
用均值填充后的数据:
姓名 数学 英语 物理
0 张三 85.0 88.0 89.5
1 李四 90.0 86.25 85.0
2 王五 86.25 82.0 90.0
3 赵六 78.0 85.0 88.0
4 钱七 92.0 90.0 95.0
Code explanation:
- Use
df['列名'].mean()to calculate the mean of each column. - Using the calculated mean as the fill value preserves the overall distribution characteristics of the data while solving the missing value problem.
Example 4: Use Forward Fill (ffill) and Backward Fill (bfill)
Forward fill fills missing values with the previous valid value, and backward fill fills missing values with the next valid value.
Example
import numpy as np
# Create a DataFrame containing missing values
data = {
'Date': ['2024-01-01', '2024-01-02', '2024-01-03', '2024-01-04', '2024-01-05'],
'Temperature': [20, np.nan, 22, np.nan, 25],
'Humidity': [60, 65, np.nan, 70, 75]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Forward fill: fill missing values with the previous value
df_ffill = df.fillna(method='ffill')
print("Data after forward fill (ffill):")
print(df_ffill)
print("=" * 50)
# Backward fill: fill missing values with the next value
df_bfill = df.fillna(method='bfill')
print("Data after backward fill (bfill):")
print(df_bfill)
Expected output:
原始数据:
日期 温度 湿度
0 2024-01-01 20.0 60.0
1 2024-01-02 NaN 65.0
2 2024-01-03 22.0 NaN
3 2024-01-04 NaN 70.0
4 2024-01-05 25.0 75.0
==================================================
向前填充(ffill)后的数据:
日期 温度 湿度
0 2024-01-01 20.0 60.0
1 2024-01-02 20.0 65.0 # 用前一个值20填充
2 2024-01-03 22.0 65.0 # 用前一个值65填充
3 2024-01-04 22.0 70.0 # 用前一个值22填充
4 2024-01-05 25.0 75.0
==================================================
向后填充(bfill)后的数据:
日期 温度 湿度
0 2024-01-01 20.0 60.0
1 2024-01-02 22.0 65.0 # 用后一个值22填充
2 2024-01-03 22.0 70.0 # 用后一个值70填充
填充
3 2024-01-04 25.0 70.0 # 用后一个值25填充
4 2024-01-05 25.0 75.0
Code explanation:
- Forward fill (
ffill): the temperature in row 1 is missing, fill it with 20 from row 0; the temperature in row 3 is missing, fill it with 22 from row 2. - Backward fill (
bfill): the temperature in row 1 is missing, fill it with 22 from row 2; the temperature in row 3 is missing, fill it with 25 from row 4. - This method is suitable for time series data, such as stock prices, temperature records, etc.
Example 5: Use the limit Parameter to Limit the Number of Fills
limitThe limit parameter can limit the maximum number of consecutive fills.
Example
import numpy as np
# Create a DataFrame containing multiple consecutive missing values
data = {
'A': [1, np.nan, np.nan, np.nan, 5],
'B': [1, 2, np.nan, np.nan, 5]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Forward fill, but limit to filling only 1 consecutive missing value at a time
df_filled = df.fillna(method='ffill', limit=1)
print(Data after limiting consecutive filling count to 1:)
print(df_filled)
Expected output:
原始数据:
A B
0 1.0 1.0
1 NaN 2.0
2 NaN NaN
3 NaN NaN
4 5.0 5.0
==================================================
限制连续填充数量为1后的数据:
A B
0 1.0 1.0
1 1.0 2.0 # 填充1个
2 NaN 2.0 # 达到limit,不再填充
3 NaN NaN # 继续不填充
4 5.0 5.0
Code explanation:
- For column A, the missing value in row 1 is filled with 1, but rows 2 and 3 are no longer filled because limit=1 is reached.
- For column B, the missing value in row 2 is filled with 2, and row 3 is no longer filled because limit=1 is reached.
Example 6: Fill Missing Values with Interpolation
Pandas also providesinterpolate()method for interpolation filling, suitable for numeric data.
Example
import numpy as np
# Create a DataFrame containing missing values
data = {
'x': [1, 2, 3, 4, 5],
'y': [10, np.nan, 30, np.nan, 50]
}
df = pd.DataFrame(data)
print(Raw data:)
print(df)
print("=" * 50)
# Use linear interpolation to fill
df_interpolated = df.copy()
df_interpolated['y'] = df_interpolated['y'].interpolate(method='linear')
print(Data after linear interpolation filling:)
print(df_interpolated)
Expected output:
原始数据: x y 0 1 10.0 1 2 NaN 2 3 30.0 3 4 NaN 4 5 50.0 ================================================== 线性插值填充后的数据: x y 0 1 10.0 1 2 20.0 # 插值:10 + (30-10)/(3-1)*(2-1) = 20 2 3 30.0 3 4 40.0 # 插值:30 + (50-30)/(5-3)*(4-3) = 40 4 5 50.0
Code explanation:
- Linear interpolation calculates missing values based on the known values before and after.
- Row 1: 10 + (30-10)/(3-1)*(2-1) = 20.
- Row 3: 30 + (50-30)/(5-3)*(4-3) = 40.
- The interpolation method is suitable for data where values change linearly.
Notes
fillna()By default, the original DataFrame is not modified. If you want to modify in place, useinplace=Trueparameter.- Note
methodThe parameter may be deprecated in future versions of Pandas. It is recommended to useffill()andbfill()method instead. - When using mean filling, if there are extreme values (outliers) in the data, consider using median filling, which is more robust.
- For categorical data (such as department, position), it is recommended to fill with the most frequent value (mode) or a clear category (such as "Unknown").
- Before filling missing values, it is recommended to analyze the causes and patterns of missingness first, and choose the most appropriate filling method.
Other extensions
Pandas Common Functions