Pandas apply / map / applymap
apply, map, and applymap are the three major functions in Pandas for data transformation. They can perform flexible element-wise or batch operations on DataFrame or Series.
Series.map
mapIs a Series method, used to transform each element in a Series.
Basic Usage
Example
# Create Series
s = pd.Series([1, 2, 3, 4, 5])
print("Original data:")
print(s)
print()
# Use a function
print("Each element * 2:")
print(s.map(lambda x: x * 2))
print()
# Use dictionary mapping
mapping = {1: "A", 2: "B", 3: "C", 4: "D", 5: "E"}
print("Using dictionary mapping:")
print(s.map(mapping))
print()
# Use Series mapping
mapping_series = pd.Series(["A", "B", "C", "D", "E"], index=[1, 2, 3, 4, 5])
print("Using Series mapping:")
print(s.map(mapping_series))
Handling Missing Values
Example
import numpy as np
s = pd.Series([1, 2, np.nan, 4, 5])
print("Contains NaN:")
print(s)
print()
# map skips NaN by default
print("map handling (skips NaN):")
print(s.map(lambda x: x * 2 if pd.notna(x) else -1))
DataFrame.applymap
applymapIs a DataFrame method that applies a function to each element one by one (Note: Pandas 2.0+ recommends usingDataFrame.mapinstead).
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, 3],
"B": [4, 5, 6],
"C": [7, 8, 9]
})
print("Original data:")
print(df)
print()
# Multiply each element by 2
print("Each element * 2:")
print(df.applymap(lambda x: x * 2))
print()
# Keep 2 decimal places
print("Keep 2 decimal places:")
print(df.applymap(lambda x: round(x, 2)))
applymap performs element-wise operations and can be slow for large data. If you only need to operate on numeric columns, consider using vectorized operations or apply with the axis parameter.
DataFrame.apply
applyIs the most flexible method and can apply functions along an axis.
Apply by Column
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, 3, 4, 5],
"B": [10, 20, 30, 40, 50],
"C": [100, 200, 300, 400, 500]
})
print("Original data:")
print(df)
print()
# Default axis=0, apply by column
print("Column sum:")
print(df.apply(sum))
print()
print("Column max:")
print(df.apply(max))
Apply by Row
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, 3],
"B": [10, 20, 30],
"C": [100, 200, 300]
})
print("Original data:")
print(df)
print()
# axis=1, apply by row
print("Row sum:")
print(df.apply(sum, axis=1))
print()
# Max - min per row
print("Row range:")
print(df.apply(lambda x: x.max() - x.min(), axis=1))
Using aggfunc for Aggregation
Example
import numpy as np
df = pd.DataFrame({
"A": [1, 2, 3],
"B": [10, 20, 30]
})
# Apply multiple functions at once
print("Sum and mean simultaneously:")
print(df.apply([sum, np.mean]))
print()
# Return multiple values
result = df.apply(lambda x: pd.Series({
"sum": x.sum(),
"mean": x.mean(),
"max": x.max()
}, index=["sum", "mean", "max"]))
print("Returning multiple values:")
print(result)
Series.apply
Series can also use apply, which is similar to map but more flexible.
Example
import numpy as np
s = pd.Series([1, 4, 9, 16, 25])
print("Original data:")
print(s)
print()
# Square root
print("Square root:")
print(s.apply(np.sqrt))
print()
# Conditional return
print("Conditional judgment:")
print(s.apply(lambda x: "large" if x > 10 else "small"))
Performance Comparison
Example
import numpy as np
import time
# Create large data
n = 100000
s = pd.Series(np.random.randn(n))
# Test map vs apply
func = lambda x: x * 2 + 1
start = time.time()
result1 = s.map(func)
map_time = time.time() - start
start = time.time()
result2 = s.apply(func)
apply_time = time.time() - start
# Vectorized (fastest)
start = time.time()
result3 = s * 2 + 1
vec_time = time.time() - start
print(f"map time: {map_time:.4f}s")
print(f"apply time: {apply_time:.4f}s")
print(f"vectorized time: {vec_time:.4f}s")
print("\nConclusion: prioritize vectorized operations for best performance")
Practical: Data Transformation
Example
import numpy as np
# Create a sample DataFrame
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
"Age": [25, 30, 28, 35],
"Salary": [12000, 15000, 11000, 18000],
"Department": ["Technology", "Sales", "Technology", "Operations"]
})
print("Original data:")
print(df)
print()
# Use apply for row-level calculation
def calculate(row):
"""Calculate annual income and after-tax salary"""
annual = row["Salary"] * 12
tax = annual * 0.1 if annual > 120000 else annual * 0.05
after_tax = annual - tax
return pd.Series({
"Annual salary": annual,
"Tax": tax,
"After tax": after_tax
})
result = df.apply(calculate, axis=1)
df_result = pd.concat([df, result], axis=1)
print("Calculation results:")
print(df_result)
Choosing Among the Three
| Method | Applicable object | Scenario | Performance |
|---|---|---|---|
map |
Series | Element-wise transformation, dictionary mapping | Fast |
applymap |
DataFrame | Element-wise transformation (non-numeric columns) | Slow |
apply |
Series/DataFrame | Row/column aggregation, custom functions | Medium |
Other ExtensionsWhen vectorized operations (directly using operators) are available, don't use apply/map; when map is available, don't use apply.