Pandas pd.get_dummies() Function
pd.get_dummies()It is a Pandas library function used forone-hot encoding of categorical variablesfunction. It converts categorical variables into binary (0/1) columns, with each category corresponding to a column.
One-hot encoding is a common technique in machine learning preprocessing because most algorithms cannot directly handle categorical data and need to convert it into numerical form.
Word Explanation: get_dummiesHere "dummy" means "dummy variable", referring to binary variables used in statistics and econometrics to represent categorical variables.
Basic Syntax and Parameters
pd.get_dummies()It is a top-level function in the Pandas library, used to convert categorical variables into one-hot encoding format.
Syntax Format
pd.get_dummies(data, prefix=None, prefix_sep='_', dummy_na=False, columns=None, drop_first=False, dtype=None)
Parameter Description
- Parameter:
data- Type: Series, DataFrame, or array-like object.
- Description: The data to be one-hot encoded. Usually a Series or DataFrame containing categorical variables.
- Parameter:
prefix- Type: String, list of strings, or dictionary.
- Description: The prefix for the generated new column names. If not specified, the original column names are used as the prefix.
- Parameter:
prefix_sep- Type: String.
- Description: The separator between the prefix and the category name. Defaults to underscore
'_'。
- Parameter:
dummy_na- Type: Boolean.
- Description: If
True, it also creates a separate column for missing values (NaN). Defaults toFalse。
- Parameter:
columns- Type: List or None.
- Description: The column names to encode. If not specified, all columns of type object, category, or boolean are encoded.
- Parameter:
drop_first- Type: Boolean.
- Description: If
True, it will remove the first column of each categorical variable to avoid multicollinearity. In models such as logistic regression, it can avoid the dummy variable trap. Defaults toFalse。
Function Description
- Return Value: Returns a DataFrame where each column corresponds to a category and the values are 0 or 1.
- Effect: Converts categorical variables into numerical binary columns, making it easier for machine learning algorithms to process.
Examples
Let's, through a series of examples from simple to complex, thoroughly masterpd.get_dummies()the usage of pd.get_dummies().
Example 1: Basic Usage - One-Hot Encoding a Series
Example
# 1. Create a Series containing categorical variables
colors = pd.Series(['red', 'blue', 'green', 'red', 'green', 'blue'])
print("=== Original Series ===")
print(colors)
# 2. Use pd.get_dummies() for one-hot encoding
result = pd.get_dummies(colors)
print("\n"=== pd.get_dummies() one-hot encoding result ===")
print(result)
Expected output:
=== 原始 Series ===
0 red
1 blue
2 green
3 red
4 green
5 blue
dtype: object
=== pd.get_dummies() 独热编码结果 ===
blue green red
0 False False True
1 True False False
2 False True False
3 False False True
4 False True False
5 True False False
Code explanation:
- The original Series contains three colors: red, blue, green.
- After one-hot encoding, each color becomes a column, using
True/Falseto indicate whether the row belongs to that category. - Each row has exactly one
True, corresponding to the original color value.
Example 2: Encoding Specific Columns of a DataFrame
In data analysis, it is often only necessary to encode specific categorical columns while leaving numerical columns unchanged.
Example
# 1. Create a DataFrame containing numerical and categorical variables
df = pd.DataFrame({
'name': ['Alice', 'Bob', 'Charlie', 'Diana'],
'age': [25, 30, 35, 28],
'city': ['Beijing', 'Shanghai', 'Beijing', 'Guangzhou'],
'department': ['Sales', 'Engineering', 'Sales', 'HR']
})
print("=== Original DataFrame ===")
print(df)
# 2. Perform one-hot encoding on the specified categorical columns
result = pd.get_dummies(df, columns=['city', 'department'])
print("\n"=== One-hot encoding for city and department columns ===")
print(result)
Expected output:
=== 原始 DataFrame ===
name age city department
0 Alice 25 Beijing Sales
1 Bob 30 Shanghai Engineering
2 Charlie 35 Beijing Sales
3 Diana 28 Guangzhou HR
=== 对 city 和 department 列进行独热编码 ===
name age city_Beijing city_Guangzhou city_Shanghai department_Engineering department_HR department_Sales
0 Alice 25 True False False False True
1 Bob 30 False False True True False
2 Charlie 35 True False False False True
3 Diana 28 False True False False False
Code explanation:
- Use the
columnsparameter to specify that only thecityanddepartmentcolumns are encoded. - Numerical columns
ageand text columnsnameremain unchanged. - The newly generated column names use the default separator underscore, such as
city_Beijing。
Example 3: Custom Prefix and Separator
You can use theprefixandprefix_sepparameter to customize the names of new columns.
# 1. Create a DataFrame
df = pd.DataFrame({
'color': ['red', 'blue', 'green', 'red'],
'size': ['S', 'M', 'L', 'XL']
})
print("=== Original DataFrame ===")
print(df)
# 2. Use the prefix parameter to customize the prefix
result_prefix = pd.get_dummies(df, prefix=['color', 'size'])
print("\n"=== Using custom prefix ===")
print(result_prefix)
# 3. Use prefix_sep to customize the separator
result_sep = pd.get_dummies(df, prefix=['color', 'size'], prefix_sep='-')
print("\n"=== Using custom separator '-' ===")
print(result_sep)
Expected output:
=== 原始 DataFrame === color size 0 red S 1 blue M 2 green L 3 red XL === 使用自定义前缀 === color_blue color_green color_red size_L size_M size_S size_XL 0 False False True False False True False 1 True False False False True False False 2 False True False True False False False 3 False False True False False False True === 使用自定义分隔符 '-' === color-blue color-green color-red size-L ...
Code explanation:
prefix=['color', 'size']Specify different prefixes for different columns.prefix_sep='-'Change the default underscore to a hyphen, and the new column names becomecolor-redthis form.
Example 4: Handling Missing Values and the drop_first Parameter
Example
import numpy as np
# 1. Data containing missing values
df = pd.DataFrame({
'color': ['red', 'blue', np.nan, 'red', 'green'],
'size': ['S', 'M', 'L', np.nan, 'XL']
})
print("=== DataFrame containing missing values ===")
print(df)
# 2. By default, missing values are not handled
result_default = pd.get_dummies(df, columns=['color'])
print("\n"=== Default: missing values not handled ===")
print(result_default)
# 3. dummy_na=True creates a separate column for missing values
result_na = pd.get_dummies(df, columns=['color'], dummy_na=True)
print("\n"=== dummy_na=True creates a column for missing values ===")
print(result_na)
# 4. drop_first=True removes the first column to avoid multicollinearity
result_drop = pd.get_dummies(df, columns=['size'], drop_first=True)
print("\n"=== drop_first=True removes the first column ===")
print(result_drop)
Expected output:
=== 包含缺失值的 DataFrame === color size 0 red S 1 blue M 2 NaN L 3 red NaN 4 green XL === 默认不处理缺失值 === size color_blue color_green color_red 0 S False False True 1 M True False False 2 L False False False 3 S False False True 4 XL False True False === dummy_na=True 为缺失值创建列 === size color_True color_blue color_green color_red 0 S False False False True 1 M False True False False 2 L True False False False 3 S False False False True 4 XL False False True False === drop_first=True 删除第一列 === color_blue color_green color_red 0 False False True 1 True False False 2 False False False 3 False False True 4 False True False
Code explanation:
dummy_na=TrueCreate an extra column for missing values (shown as the color_True column in this example).drop_first=TrueRemoving the first column of each group (e.g., size_M is removed) can avoid multicollinearity issues in linear models.
Tip:One-hot encoding increases the data dimensionality (one column per category). If the number of categories is very large, it may lead to the curse of dimensionality. In such cases, you can consider using label encoding (
pd.factorize()) or other encoding methods.
Common Pandas Functions