Pandas df.dropna() Function
df.dropna()is a function in Pandas used to delete rows or columns that contain missing values.
) is a common problem.NaN、NoneorNULLThe function provides a simple way to clean data. You can choose to delete rows or columns containing missing values, or filter data based on the proportion of missing values.dropna()Basic Syntax and Parameters
Basic Syntax and Parameters
dropna()to call..Syntax Format
Syntax Format
DataFrame.dropna(axis=0, how='any', thresh=None, subset=None, inplace=False)
Parameter Description
| Type | Required or Not | Description | Default Value | int or str |
|---|---|---|---|---|
| axis | Optional | Specifies whether to delete rows or columns. | Indicates deleting rows with missing values;0or'index'Indicates deleting columns with missing values.1or'columns'Optional |
0 |
| how | str | Specifies the deletion condition. | Indicates deletion if there are any missing values;'any'Indicates deletion only if all values are missing.'all'Optional |
'any' |
| thresh | int | Sets the minimum number of non-missing values. If the number of non-missing values in a row (or column) is less than | , then the row (or column) is deleted.threshOptional |
None |
| subset | array-like | Specifies in which columns (or rows) to check for missing values. If | , thenaxis=0specifies column names; ifsubset, thenaxis=1specifies row indices.subsetOptional |
None |
| inplace | bool | If | , modifies the original DataFrame directly and does not return a new object; ifTrue, returns a new DataFrame, leaving the original data unchanged.FalseReturn Value Description |
False |
Return Value Description
- ), or
inplace=False(ifNoneThe returned DataFrame does not contain rows or columns with missing values.inplace=True)。 - Examples
Examples
the usage of.dropna()Example 1: Delete Rows with Missing Values
Example 1: Delete Rows with Missing Values
Example
Example
import numpy as np
'Name'
data = {
'Zhang San': ['Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi', 'age'],
# Wang Wu's age is missing: [25, 30, np.nan, 35, 28], 'Salary'
# Zhao Liu's salary is missing: [5000, 6000, 5500, np.nan, 7000], 'Department'
'Technology': ['Marketing', 'Technology', 'Marketing', 'Technology', “Raw data:”]
}
df = pd.DataFrame(data)
print(# Delete rows with missing values (default behavior))
print(df)
print("=" * 50)
"Data after removing missing values:"
df_cleaned = df.dropna()
print(Expected output:)
print(df_cleaned)
Code analysis:
原始数据:
姓名 年龄 薪资 部门
0 张三 25.0 5000.0 技术
1 李四 30.0 6000.0 市场
2 王五 NaN 5500.0 技术
3 赵六 35.0 NaN 市场
4 钱七 28.0 7000.0 技术
==================================================
删除缺失值后的数据:
姓名 年龄 薪资 部门
0 张三 25.0 5000.0 技术
1 李四 30.0 6000.0 市场
4 钱七 28.0 7000.0 技术
We created a DataFrame containing missing values, where Wang Wu's age is
- , and Zhao Liu's salary is
NaNUsingNaN。 - with default parameters, all rows containing missing values were deleted.
df.dropna()The result kept the data of Zhang San, Li Si, and Qian Qi, and deleted Wang Wu and Zhao Liu. - Example 2: Delete Columns with Missing Values
Example 2: Delete Columns with Missing Values
Example
Example
import numpy as np
'Name'
data = {
'Zhang San': ['Li Si', 'Wang Wu', 'Zhao Liu', 'age'],
'Salary': [25, 30, 35, 40],
# Multiple columns have missing values: [5000, np.nan, 5500, np.nan], 'Department'
'Marketing': [np.nan, 'Technology', 'Marketing', 'city'],
'Beijing': ['Shanghai', 'Guangzhou', 'Shenzhen', # No missing values] “Raw data:”
}
df = pd.DataFrame(data)
print(# Delete columns with missing values)
print(df)
print("=" * 50)
Data after removing columns with missing values:
df_cleaned = df.dropna(axis=1)
print(Expected output:)
print(df_cleaned)
Code analysis:
原始数据:
姓名 年龄 薪资 部门 城市
0 张三 25 5000.0 NaN 北京
1 李四 30 NaN 市场 上海
2 王五 35 5500.0 技术 广州
3 赵六 40 NaN 市场 深圳
==================================================
删除缺失值列后的数据:
姓名 年龄 城市
0 张三 25 北京
1 李四 30 上海
2 王五 35 广州
3 赵六 40 深圳
Setting
- indicates operating by column.
axis=1The Salary and Department columns contain missing values, so they are deleted. - The Name, Age, and City columns have no missing values, so they are retained.
- Example 3: Use the thresh Parameter to Control Deletion Conditions
Example 3: Use the thresh Parameter to Control Deletion Conditions
threshExample
Example
import numpy as np
# Has 2 missing values
data = {
'A': [1, 2, np.nan, 4, 5],
'B': [np.nan, 2, 3, 4, np.nan], # No missing values
'C': [1, 2, 3, 4, 5], # Has 3 missing values
'D': [np.nan, np.nan, np.nan, 4, 5] “Raw data:”
}
df = pd.DataFrame(data)
print(# Keep only rows with at least 3 non-missing values)
print(df)
print("=" * 50)
"Keep rows with at least 3 non-missing values:"
df_cleaned = df.dropna(thresh=3)
print(Expected output:)
print(df_cleaned)
Code analysis:
原始数据:
A B C D
0 1.0 NaN 1.0 NaN
1 2.0 2.0 2.0 NaN
2 NaN 3.0 3.0 NaN
3 4.0 4.0 4.0 4.0
4 5.0 NaN 5.0 5.0
==================================================
保留至少3个非缺失值的行:
A B C D
1 2.0 2.0 2.0 NaN
3 4.0 4.0 4.0 4.0
Row 0 has only 1 non-missing value (1.0), so it is deleted.
- Row 1 has 3 non-missing values (2.0, 2.0, 2.0), so it is retained.
- Row 2 has only 2 non-missing values (3.0, 3.0), so it is deleted.
- Row 3 has 4 non-missing values, so it is retained.
- Row 4 has only 3 non-missing values (5.0, 5.0, 5.0), so it is retained.
- Example 4: Use subset to Specify Columns to Check
Example 4: Use subset to Specify Columns to Check
parameter.subsetExample
Example
import numpy as np
'Name'
data = {
'Zhang San': ['Li Si', 'Wang Wu', 'Zhao Liu', 'age'],
# Li Si's age is missing: [25, np.nan, 35, 40], 'Salary'
# Wang Wu's salary is missing: [5000, 6000, np.nan, 8000], 'Department'
'Technology': ['Marketing', 'Technology', 'Marketing', “Raw data:”]
}
df = pd.DataFrame(data)
print(# Check for missing values only in the "Name" and "Department" columns)
print(df)
print("=" * 50)
'Name'
df_cleaned = df.dropna(subset=['Department', Data after checking for missing values only in Name and Department columns:])
print(# Check for missing values only in the "薪资" column)
print(df_cleaned)
print("=" * 50)
'Salary'
df_cleaned2 = df.dropna(subset=["Data after checking for missing values only in the salary column:"])
print(Expected output:)
print(df_cleaned2)
Code analysis:
原始数据:
姓名 年龄 薪资 部门
0 张三 25.0 5000.0 技术
1 李四 NaN 6000.0 市场
2 王五 35.0 NaN 技术
3 赵六 40.0 8000.0 市场
==================================================
只在姓名和部门列检查缺失值后的数据:
姓名 年龄 薪资 部门
0 张三 25.0 5000.0 技术
1 李四 NaN 6000.0 市场
2 王五 35.0 NaN 技术
3 赵六 40.0 8000.0 市场
==================================================
只在薪资列检查缺失值后的数据:
姓名 年龄 薪资 部门
0 张三 25.0 5000.0 技术
1 李四 30.0 6000.0 市场
3 赵六 40.0 8000.0 市场
The Name and Department columns have no missing values, so all rows are retained after the check.
- In the Salary column, Wang Wu's salary is
- , so row 2 is deleted.
NaNExample 5: Use how='all' to Only Delete Rows That Are All Missing Values
Example 5: Use how='all' to Only Delete Rows with All Missing Values
how='all'Example
Example
import numpy as np
“Raw data:”
data = {
'A': [1, np.nan, np.nan, 4],
'B': [2, np.nan, np.nan, 5],
'C': [3, np.nan, np.nan, 6]
}
df = pd.DataFrame(data)
print(# Delete rows where all values are missing)
print(df)
print("=" * 50)
After removing all rows with missing values:
df_cleaned = df.dropna(how='all')
print(Expected output:)
print(df_cleaned)
Code analysis:
原始数据:
A B C
0 1.0 2.0 3.0
1 NaN NaN NaN
2 NaN NaN NaN
3 4.0 5.0 6.0
==================================================
删除全部为缺失值的行后:
A B C
0 1.0 2.0 3.0
3 4.0 5.0 6.0
All values in rows 1 and 2 are
- , using
NaNdelete them.how='all'Delete them. - Rows 0 and 3 contain actual data and were retained.
Notes
dropna()By default, the original DataFrame is not modified. To modify it in place, useinplace=Trueparameter.- Use
threshWhen using the parameter, the value you set should be based on your requirements for data quality. - Before deleting data, it is recommended to first use
df.isnull().sum()to view the number of missing values in each column, so you can make more informed decisions. - For important datasets, use with caution
dropna(), because this may lead to a large amount of data loss. You may consider usingfillna()to fill missing values.
Other Extensions