Pandas df.replace() function
df.replace()It is a function in Pandas used to replace data in a DataFrame.
Data replacement is a common operation in data cleaning,replace()It can help you replace specific values with new values, and supports multiple flexible methods such as single-value replacement, multi-value replacement, and regular expression replacement. This is very useful in scenarios such as handling outliers, unifying data formats, and encoding conversion.
Basic syntax and parameters
replace()It is a member function of DataFrame, through the dot operator.to call.
Syntax format
DataFrame.replace(to_replace=None, value=None, inplace=False, limit=None, regex=False, method='pad')
Parameter description
| Parameter | Type | Required | Description | Default value |
|---|---|---|---|---|
| to_replace | str, regex, list, dict, int, float, None | Required | The value to be replaced. It can be a single value, a list of values, a dictionary, a regular expression, etc. | None |
| value | scalar, dict, list, str, regex, None | Optional | The value after replacement. Ifto_replaceit is a dictionary, this parameter can be omitted; if it is not a dictionary, this parameter is required. |
None |
| inplace | bool | Optional | If it isTrue, modify the original DataFrame directly without returning a new object; if it isFalse, return a new DataFrame, the original data remains unchanged. |
False |
| limit | int | Optional | Specify the maximum number of replacements. | None |
| regex | bool or str | Optional | If it isTrue, treatto_replaceas a regular expression. |
False |
| method | str | Optional | Whento_replaceUsed when it is a list.'pad'or'ffill'Indicates forward fill;'backfill'or'bfill'Indicates backward fill. |
'pad' |
Return value description
- Returns a new DataFrame (if
inplace=False), orNone(ifinplace=True)。 - The specified values in the returned DataFrame have been replaced.
Examples
Let's thoroughly master through a series of examplesreplace()the usage of.
Example 1: Single value replacement
Replace a specific value in the DataFrame with a new value.
Example
import numpy as np
# Create a DataFrame that contains duplicate data
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'Department': ['Technology', 'Marketing', 'Technology', 'Marketing'],
'Salary': [5000, 6000, 5500, 7000]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Replace 'technology' with 'R&D'
df_replaced = df.replace('Technology', R&D)
print("Replaced Data:")
print(df_replaced)
Expected output:
原始数据:
姓名 部门 薪资
0 张三 技术 5000
1 李四 市场 6000
2 王五 技术 5500
3 赵六 市场 7000
==================================================
替换后的数据:
姓名 部门 薪资
0 张三 研发 5000
1 李四 市场 6000
2 王五 研发 5500
3 赵六 市场 7000
Code analysis:
- There are two rows with "Technology" department in the DataFrame.
- Use
df.replace('技术', '研发')to replace all "technology" with "R&D". - This method replaces all matching values in the DataFrame.
Example 2: One-to-one replacement of multiple values
Use a dictionary to replace multiple different values at once.
Example
# Create a DataFrame containing the data to be replaced
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen'],
Level: ['A', 'B', 'A', 'C']
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Use a dictionary to perform multiple replacements
replacements = {
'Beijing': 'Beijing',
'Shanghai': 'Shanghai',
'Guangzhou': 'Guangzhou',
'Shenzhen': 'Shenzhen',
'A': 'Excellent',
'B': Good,
'C': 'Pass'
}
df_replaced = df.replace(replacements)
print("Replaced Data:")
print(df_replaced)
Expected output:
原始数据:
姓名 城市 等级
0 张三 北京 A
1 李四 上海 B
2 王五 广州 A
3 赵六 深圳 C
==================================================
替换后的数据:
姓名 城市 等级
0 张三 北京市 优秀
1 李四 上海市 良好
2 王五 广州市 优秀
3 赵六 深圳市 及格
Code analysis:
- Through the dictionary
replacements, we can specify multiple replacement rules at once. - This method is very efficient, avoiding multiple calls to
replace()。
Example 3: Replace values in a specific column
You can replace values only in specific columns without affecting other columns.
Example
# Create a DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'Department': ['Technology', 'Marketing', 'Technology', 'Marketing'],
Position: ['Technology', 'Marketing', 'Technology', 'Marketing']
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Only replace "Technology" with "R&D" in the "Department" column
df_replaced = df.replace({'Department': 'Technology'}, R&D)
print(Data after replacing only the department column:)
print(df_replaced)
print("=" * 50)
You can also use a nested dictionary to apply different replacement rules to different columns.
df_replaced2 = df.replace({'Department': {'Technology': R&D, 'Marketing': Sales}, Position: 'Technology'})
print("Data after applying different replacement rules to different columns:")
print(df_replaced2)
Expected output:
原始数据:
姓名 部门 职位
0 张三 技术 技术
1 李四 市场 市场
2 王五 技术 技术
3 赵六 市场 市场
==================================================
只替换部门列后的数据:
姓名 部门 职位
0 张三 研发 技术
1 李四 市场 市场
2 王五 研发 技术
3 赵六 市场 市场
==================================================
对不同列使用不同规则替换后的数据:
姓名 部门 职位
0 张三 研发 研发
1 李四 销售 市场
2 王五 研发 研发
3 赵六 销售 市场
Code analysis:
- Using
{'department': 'Technology'}can replace only the values in the "Department" column. - Nested dictionary
{'department': {'Technology': 'R&D'}}can specify different replacement rules for different columns.
Example 4: Replace using regular expressions
regex=TrueThe parameter allows the use of regular expressions for replacement, which is very useful when handling text with pattern matching.
Example
# Create a DataFrame that needs to be processed with regular expressions
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
Phone: ['138-0000-0000', '139-1111-1111', '137-2222-2222', '136-3333-3333'],
Remarks: [Normal, 'VIPclient', Normal, 'VIP customer']
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Use regular expressions to replace separators in phone numbers
df_replaced = df.replace(to_replace=r'(d{3})-(d{4})-(d{4})', value=r'1****3', regex=True)
print(Data after replacing phone numbers (hiding the middle 4 digits):)
print(df_replaced)
print("=" * 50)
# Replace "VIP customer" in the remark (there may be whitespace differences)
df_replaced2 = df.replace(to_replace=r'VIPs*Customers?', value='VIP', regex=True)
print(Data after unifying VIP remarks:)
print(df_replaced2)
Expected output:
原始数据:
姓名 电话 备注
0 张三 138-0000-0000 正常
1 李四 139-1111-1111 VIP客户
2 王五 137-2222-2222 正常
3 赵六 136-3333-3333 VIP 客户
==================================================
替换电话号码(隐藏中间4位)后的数据:
姓名 电话 备注
0 张三 138****0000 正常
1 李四 139****1111 VIP客户
2 王五 137****2222 正常
3 赵六 136****3333 VIP 客户
==================================================
统一VIP备注后的数据:
姓名 电话 备注
0 张三 138-0000-0000 正常
1 李四 139-1111-1111 VIP
2 王五 137-2222-2222 正常
3 赵六 136-3333-3333 VIP
Code analysis:
- The first example uses a regular expression
(d{3})-(d{4})-(d{4})to match the phone number format and replace it with1****3, hiding the middle 4 digits. - The second example uses
VIPs*客户?to match both "VIPclient" and "VIP client" formats, and uniformly replace them with "VIP".
Example 5: Replace missing values NaN
replace()It can also be used to replace missing valuesNaN, which is a common operation in data cleaning.
Example
import numpy as np
# Create a DataFrame with missing values
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'Age': [25, np.nan, 35, np.nan],
'Salary': [5000, 6000, np.nan, 8000]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Replace NaN with 0
df_replaced = df.replace(np.nan, 0)
print("Data after replacing NaN with 0:")
print(df_replaced)
print("=" * 50)
# Replace NaN with other values, such as "Unknown"
df_replaced2 = df.replace(np.nan, 'Unknown')
print("Data after replacing NaN with 'Unknown':")
print(df_replaced2)
Expected output:
原始数据:
姓名 年龄 薪资
0 张三 25.0 5000.0
1 李四 NaN 6000.0
2 王五 35.0 NaN
3 赵六 NaN 8000.0
==================================================
将 NaN 替换为 0 后的数据:
姓名 年龄 薪资
0 张三 25.0 5000
1 李四 0.0 6000
2 王五 35.0 0
3 赵六 0.0 8000
==================================================
将 NaN 替换为'未知'后的数据:
姓名 年龄 薪资
0 张三 25.0 5000.0
1 李四 未知 6000.0
2 王五 35.0 未知
3 赵六 未知 8000.0
Code explanation:
np.nanrepresents missing values. You can usereplace(np.nan, value)to replace them.- Choose an appropriate value to replace missing values based on the data type and business requirements.
Example 6: Numeric replacement
You can replace numeric data to perform data standardization or outlier handling.
Example
# Create a DataFrame containing outliers
data = {
'Student': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
'Score': [85, 150, 78, 92, -10] # 150 and -10 are outliers
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Replace outliers with values in a reasonable range
df_replaced = df.replace({150: 100, -10: 0})
print("Data after replacing outliers:")
print(df_replaced)
print("=" * 50)
# Use conditional logic to replace multiple values
df_replaced2 = df.replace(to_replace=[150, -10], value=[100, 0])
print("Using lists to replace multiple values:")
print(df_replaced2)
Expected output:
原始数据:
学生 成绩
0 张三 85
1 李四 150
2 王五 78
3 赵六 92
4 钱七 -10
==================================================
将异常值替换后的数据:
学生 成绩
0 张三 85
1 李四 100
2 王五 78
3 赵六 92
4 钱七 0
==================================================
使用列表替换多个值:
学生 成绩
0 张三 85
1 李四 100
2 王五 78
3 赵六 92
4 钱七 0
Code explanation:
- dictionary
{150: 100, -10: 0}Replace outlier 150 with 100, and -10 with 0. - list
[150, -10]and[100, 0]Specify the values to be replaced and the replacement values respectively, corresponding in order.
Notes
replace()By default, the original DataFrame is not modified. If you want to modify it in place, useinplace=Trueparameter.- When using regular expressions, ensure the regex syntax is correct; complex regular expressions may cause unexpected results.
- Replacement is based on value matching and does not change the data type.
- Pay attention to case sensitivity; "Technology" and "Technology" are different values.
- Before data replacement, it is recommended to back up the original data for comparison and traceability.
Other Extensions