Pandas Missing Value Handling
Real data often contains missing values (NaN). Pandas provides rich functions to handle missing data. This section details the use of fillna, dropna, interpolate and other methods.
Representation of Missing Values
In Pandas,NaN(Not a Number) represents missing values, which comes from the NumPy library.
Example
import numpy as np
# Create data containing missing values
s = pd.Series([1, 2, np.nan, 4, 5])
print("Series containing NaN:")
print(s)
print(f"Has missing values: {s.isna().any()}")
print()
# Missing values in DataFrame
df = pd.DataFrame({
"A": [1, 2, np.nan, 4],
"B": [np.nan, 2, 3, 4],
"C": [1, 2, 3, np.nan]
})
print("DataFrame containing missing values:")
print(df)
print()
# Detect missing values
print("Missing value positions:")
print(df.isna())
print()
# Count missing values per column
print("Number of missing values per column:")
print(df.isna().sum())
In Pandas,
NaN、None、pandas.NAare all recognized as missing values. Usingisna()orisnull()can uniformly detect these missing values.
dropna: Delete Missing Values
Delete rows
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, np.nan, 4],
"B": [1, np.nan, 3, 4],
"C": [1, 2, 3, np.nan]
})
print("Original data:")
print(df)
print()
# Delete rows containing missing values (default)
print("Delete rows with missing values:")
print(df.dropna())
print()
# how='all': Delete only if all values are missing
print("Only rows that are all missing values are deleted:")
print(df.dropna(how="all"))
print()
# thresh: Keep at least N non-missing values
print("Keep rows with at least 2 non-missing values:")
print(df.dropna(thresh=2))
Delete columns
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, np.nan, 4],
"B": [np.nan, np.nan, np.nan, np.nan], # All missing
"C": [1, 2, 3, 4]
})
print("Original data:")
print(df)
print()
# Delete columns containing missing values
print("Delete columns with missing values:")
print(df.dropna(axis=1))
print()
# how='all': Delete columns that are all missing values
print("Delete columns that are all missing values:")
print(df.dropna(axis=1, how="all"))
fillna: Fill Missing Values
Fixed value filling
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, np.nan, 4, 5],
"B": [np.nan, 2, 3, np.nan, 5],
"C": [1, 2, 3, 4, np.nan]
})
print("Original data:")
print(df)
print()
# Fill with 0
print("Fill with 0:")
print(df.fillna(0))
print()
# Fill different columns with specified values
print("Different columns filled with different values:")
print(df.fillna({"A": 0, "B": 99, "C": -1}))
print()
# Fill with previous value (forward fill)
print("Forward fill:")
print(df.fillna(method="ffill"))
print()
# Fill with next value (backward fill)
print("Backward fill:")
print(df.fillna(method="bfill"))
Statistical value filling
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, np.nan, 4, 5, 6],
"B": [10, np.nan, 30, np.nan, 50, 60]
})
print("Original data:")
print(df)
print()
# Fill with mean
print("Fill with mean:")
print(df.fillna(df.mean()))
print()
# Fill with median
print("Fill with median:")
print(df.fillna(df.median()))
print()
# Fill by column respectively
print("Column A uses mean, column B uses 0:")
print(df.fillna({"A": df["A"].mean(), "B": 0}))
Forward fill and backward fill are particularly useful in time series data and can maintain data continuity. Which method to choose depends on the business meaning of the data.
interpolate: Interpolation filling
Interpolation is a smarter filling method that can infer missing values from adjacent data.
Linear interpolation
Example
import numpy as np
s = pd.Series([1, 2, np.nan, 4, 5, np.nan, 7])
print("Original Series:")
print(s)
print()
# Linear interpolation (default)
print("Linear interpolation:")
print(s.interpolate())
print()
# Specify interpolation method
print("time weighted interpolation (time series):")
s2 = pd.Series([1, np.nan, np.nan, 4], index=pd.date_range("2024-01-01", periods=4, freq="D"))
print(s2.interpolate(method="time"))
print()
# Fill with nearest value
print("Fill with nearest value:")
print(s.interpolate(method="nearest"))
DataFrame interpolation
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, np.nan, 4, 5],
"B": [10, np.nan, 30, 40, np.nan]
})
print("Original data:")
print(df)
print()
# Interpolate on the original DataFrame
df_interpolated = df.interpolate(method="linear")
print("After linear interpolation:")
print(df_interpolated)
print()
# Limit interpolation range
print("Only fill the first and last of consecutive missing values (max 1):")
print(df.interpolate(limit=1))
Practical: Handling Real Data
Example
import numpy as np
# Simulate real business data
np.random.seed(42)
n = 20
df = pd.DataFrame({
"Date": pd.date_range("2024-01-01", periods=n),
"Sales amount": np.random.choice([100, 200, np.nan, 300, np.nan], n),
"Customer count": np.random.choice([10, np.nan, 20, 30], n),
"Conversion rate": np.random.choice([0.05, 0.1, np.nan, 0.15], n)
})
print("Missing data in original data:")
print(df.isna().sum())
print()
# Handling strategy:
# 1. Fill sales amount with the mean of preceding and following values (business allows fluctuations)
df["Sales amount"] = df["Sales amount"].interpolate()
# 2. Fill customer count with mean
df["Customer count"] = df["Customer count"].fillna(df["Customer count"].mean())
# 3. Fill conversion rate with 0 (indicating no conversion)
df["Conversion rate"] = df["Conversion rate"].fillna(0)
print("Processed data:")
print(df)
print()
print("Missing data after processing:")
print(df.isna().sum())
Common Questions
1. Does fillna create a new object or modify in place?
fillnaBy default it returns a new object. Useinplace=Truecan modify in place.
2. Still missing values after interpolation
If missing values are at the beginning or end of the sequence, interpolation cannot fill them, and additional handling is needed.
3. Distinguish between NaN and empty strings
Empty string""is not a missing value, you need to usereplace("", np.nan)to convert.
Other extensionsWhen choosing a filling method, consider the business meaning: forward filling is suitable for time series, backward filling for static data, mean for numeric data, and interpolation for continuous data.