)}"

Pandas 常用函数Pandas Common Functions


pd.factorize()Is a Pandas library function used forencoding categorical variablesThe function converts categorical data into integer codes and returns a list of unique values.

It is a common technique in machine learning preprocessing, converting text categories into numeric values, making it easier for algorithms to process. Compared with one-hot encoding (pd.get_dummies()), factorization encoding does not increase the data dimension.

Word Definition: factorizeMeans "factorization", here referring to converting categorical variables into numeric factors (integer codes).


Basic Syntax and Parameters

pd.factorize()It is a top-level function of the Pandas library, used to encode categorical variables as integers.

Syntax Format

pd.factorize(values, sort=False, na_sentinel=-1, size_hint=None)

Parameter Description

  • Parameter: values
    • Type: Array-like object, such as Series, list, array, etc.
    • Description: The categorical data to be encoded.
  • Parameter: sort
    • Type: Boolean value.
    • If it isTrue, sort the unique values before encoding. Default isFalse(encode in order of appearance).
  • Parameter: na_sentinel
    • Type: Integer.
    • Integer used to mark missing values. Default is-1. If set toNone, missing values will be kept as NaN.
  • Function Description

    • Return Value: Returns a tuple (codes, uniques), where codes is the encoded integer array, and uniques is the array of unique values.
    • Effect: Maps categorical variables to an integer sequence starting from 0, with each unique value corresponding to an integer.

    Examples

    Let's thoroughly master through a series of examples from simple to complexpd.factorize()its usage.

    Example 1: Basic Usage - Encoding Categorical Variables

    Example

    import pandas as pd
    import numpy as np

    # 1. Create categorical data
    colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow'])

    print("=== Original categorical data ===")
    print(colors)

    # 2. Use pd.factorize() for encoding
    codes, uniques = pd.factorize(colors)

    print("\n=== pd.factorize() encoding result ===")
    print(fEncoding: {codes})  # Encoded integer
    print(fUnique values: {uniques})  # uniqueValueList
    print(f"Type: {type(codes)}")

    Expected output:

    === 原始分类数据 ===
    0       red
    1      blue
    2     green
    3       red
    4      blue
    5     yellow
    原始数据中 'red' 出现 2 次,'blue' 出现 2 次,'green' 和 'yellow' 各出现 1 次。
    
    按出现顺序分配编码:
    - 'red' -> 0
    - 'blue' -> 1
    - 'green' -> 2
    - 'yellow' -> 3
    === 编码: [0 1 2 0 1 3]
    === 唯一值数组: ['red' 'blue' 'green' 'yellow']
    

    Code explanation:

    1. The returned codes is an integer array, where each original value is replaced with the corresponding integer code.
    2. The uniques array stores all unique values, arranged in the corresponding order of encoding.
    3. By default, codes are assigned in order of appearance (the first appearing value is encoded as 0).

    Example 2: sort Parameter - Sorting Unique Values Before Encoding

    Usingsort=Trueyou can sort the unique values before encoding, making the results more predictable.

    Example

    import pandas as pd
    import numpy as np

    # Create categorical data
    colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow'])

    print(=== Raw Data ===)
    print(colors)

    # 1. No sorting (default) - encode in order of appearance
    codes_unsorted, uniques_unsorted = pd.factorize(colors, sort=False)
    print("\n=== sort=False (default) ===")
    print(f"Codes: {codes_unsorted}")
    print(fUnique values: {uniques_unsorted})

    # 2. Encode after sorting - in alphabetical order
    codes_sorted, uniques_sorted = pd.factorize(colors, sort=True)
    print("\n=== sort=True (in alphabetical order) ===)
    print(fCodes: {codes_sorted})
    print(f"Unique values: {uniques_sorted}")

    # 3. Sort numeric values
    numbers = pd.Series([5, 2, 8, 5, 3, 2, 9])
    codes_num, uniques_num = pd.factorize(numbers, sort=True)
    print("\n=== Numeric Sorting Encoding ===)
    print(f"Codes: {codes_num}")
    print(fUnique values: {uniques_num})

    Expected output:

    === 原始数据 ===
    0       red
    1      blue
    2      green
    3       red
    4      blue
    5      yellow
    
    === sort=False (默认) ===
    编码: [0 1 2 0 1 3]
    唯一值: ['red' 'blue' 'green' 'yellow']
    
    === sort=True (按字母顺序) ===
    编码: [2 0 1 2 0 3]
    唯一值: ['blue' 'green', 'red', 'yellow']
    按字母顺序排列:blue, green, red, yellow
    - blue -> 0
    - green -> 1
    - red -> 2
    - yellow -> 3
    
    === 数值排序编码 ===
    编码: [2 0 3 2 1 0 4]
    唯一值: [2, 3, 5, 8, 9]
    按数值排序:2, 3, 5, 8, 9
    - 2 -> 0
    - 3 -> 1
    - 5 -> 2
    - 8 -> 3
    - 9 -> 4
    

    Code explanation:

    • sort=TrueAfter sorting the unique values and then encoding, the coding order is deterministic and predictable.
    • This is useful when reproducible results are required.

    Example 3: Handling Missing Values

    na_sentinelThe parameter is used to control the handling of missing values.

    Example

    import pandas as pd
    import numpy as np

    Data containing missing values
    data = pd.Series(['a', 'b', np.nan, 'a', None, 'c', 'b'])

    print("=== Data with missing values ===")
    print(data)

    # 2. Default behavior (missing value → -1)
    codes_default, uniques_default = pd.factorize(data)
    print("\n=== Default handling (missing values = -1) ===)
    print(fEncoding: {codes_default})
    print(f"Unique values: {uniques_default}")

    # 3. Custom missing value encoding (simulating na_sentinel=-999)
    codes_custom = np.where(codes_default == -1, -999, codes_default)
    print("\n=== Custom missing value encoding (-999) ===)
    print(fCodes: {codes_custom})

    # 4. Keep NaN (simulate na_sentinel=None)
    codes_nan = codes_default.astype('float')
    codes_nan[codes_nan == -1] = np.nan

    print("\n=== Keep NaN ===)
    print(f"Codes: {codes_nan}")
    print(f"Encoding type: {type(codes_nan)

    Expected output:

    === 包含缺失值的数据 ===
    0       a
    1       b
    2     NaN
    3       a
    4    None
    5       c
    6       b
    dtype: object
    
    === 默认处理(缺失值 = -1)===
    编码: [ 0  1 -1  0 -1  2  1]
    唯一值: Index(['a', 'b', 'c'], dtype='object')
    
    === 自定义缺失值编码(-999)===
    编码: [   0    1 -999    0 -999    2    1]
    
    === 保持 NaN ===
    编码: [ 0.  1. nan  0. nan  2.  1.]
    编码类型: <class 'numpy.float64'>

    Example 4: Application in Machine Learning

    Factorization encoding is very practical in machine learning preprocessing, especially for high-cardinality categorical variables.

    Example

    import pandas as pd

    Create a DataFrame containing categorical variables
    df = pd.DataFrame({
        'name': ['Alice', 'Bob', 'Charlie', 'Diana', 'Eve', 'Frank'],
        'city': ['Beijing', 'Shanghai', 'Beijing', 'Guangzhou', 'Shanghai', 'Beijing'],
        'department': ['Sales', 'Engineering', 'Sales', 'HR', 'Engineering', 'Sales']
    })

    print("=== Original DataFrame ===")
    print(df)

    # 2. Factorize categorical columns
    # Create a mapping dictionary
    city_codes, city_uniques = pd.factorize(df['city'])
    dept_codes, dept_uniques = pd.factorize(df['department'])

    print("\n=== Factorization Encoding Result ===)
    print(f"city codes: {city_codes.tolist()}")
    print(f"city unique values: {city_uniques}")
    print(f"department code: {dept_codes.tolist()}")
    print(f"department unique values: {dept_uniques}")

    # 3. Create the encoded DataFrame
    df_encoded = df.copy()
    df_encoded['city_encoded'] = city_codes
    df_encoded['dept_encoded'] = dept_codes

    print("\n=== Encoded DataFrame ===)
    print(df_encoded)

    # 4. Use a mapping dictionary to encode new data
    print("\n=== Mapping dictionary ===)
    city_mapping = dict(zip(city_uniques, range(len(city_uniques))))
    print(fcity mapping: {city_mapping})

    # Encode new data
    new_cities = pd.Series(['Beijing', 'Shanghai', 'Shenzhen'])
    new_codes = new_cities.map(city_mapping)
    print(f"\nNew data encoding: {new_codes.tolist()}")

    Expected output:

    === 原始 DataFrame ===
        name       city department
    0  Alice    Beijing      Sales
    1    Bob   Shanghai  Engineering
    2  Charlie    Beijing      Sales
    3  Diana  Guangzhou        HR
    4   eve  Shanghai  Engineering
    5   Frank    Beijing      Sales
    
    === 因子化编码结果 ===
    city 编码: [0 1 0 2 1 0]
    city 唯一值: ['Beijing' 'Shanghai' 'Guangzhou']
    department 编码: [0 1 0 2 1 0]
    department 唯一值: ['Sales' 'Engineering' 'HR']
    
    === 编码后的 DataFrame ===
        name       city department  city_encoded  dept_encoded
    0  Alice    Beijing      Sales             0            0
    1    Bob   Shanghai  Engineering           1            1
    2  Charlie    Beijing      Sales            0            0
    3  Diana  Guangzhou        HR               2            2
    4    Eve   Shanghai  Engineering           1            1
    5   Frank    Beijing  Sales               0            0
    
    === 映射字典 ===
    city 映射: {'Beijing': 0, 'Shanghai': 1, 'Guangzhou': 2}
    

    Code explanation:

    • pd.factorize()Returns a tuple that can be directly unpacked to obtain the codes and unique values.
    • After creating a mapping dictionary, the same encoding can be applied to new data, facilitating machine learning model inference.
    • Factorization encoding does not increase the data dimension, making it suitable for handling high-cardinality categorical variables.

    Tip: pd.factorize()Suitable for handling ordered or unordered categorical variables. If one-hot encoding is needed (for algorithms such as logistic regression), please usepd.get_dummies()function.

    Pandas 常用函数Pandas Common Functions

    Other Extensions