Pandas pd.unique() Function

Pandas 通用函数Common Pandas Functions


pd.unique()is used in the Pandas library toget unique values from an arrayfunction. It returns all non-duplicate values from the input array, removing duplicates.

This is a common operation in data analysis, such as counting how many distinct values are in a column, or getting all categories of a categorical variable.

Word Definition: uniqueIt means "unique, distinctive"; here it refers to returning non-duplicate values from the array.


Basic Syntax and Parameters

pd.unique()is a top-level function of the Pandas library, used to extract unique values from an array.

Syntax Format

pd.unique(values)

Parameter Description

  • Parameter: values
    • Type: array-like object, such as a list, Series, one-dimensional array, etc.
    • Description: The input data from which to extract unique values. It can be any one-dimensional array structure.

Function Description

  • Return Value: Returns an ndarray (NumPy array) containing all unique values.
  • Effect: Removes duplicate values from the input data, keeping each value only once.

Examples

Let's use a series of examples from simple to complex to thoroughly masterpd.unique()its usage.

Example 1: Basic Usage - Extracting Unique Values from a Series

Example

import pandas as pd
import numpy as np

# 1. Create a Series containing duplicate values
colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow', 'red'])

print(=== Original Series ===)
print(colors)

# 2. Use pd.unique() to get unique values
unique_values = pd.unique(colors)
print("\n=== pd.unique() Unique Values ===")
print(unique_values)
print(f"\nNumber of unique values: {len(unique_values)}")

Expected output:

=== 原始 Series ===
0       red
1      blue
2     green
3       red
4      blue
5    yellow
6       red
dtype: object

=== pd.unique() 唯一值 ===
['red' 'blue' 'green' 'yellow']
唯一值数量: 4

Code Explanation:

  • The original Series has 7 elements, but only 4 unique values.
  • pd.unique()Returns a NumPy array containing all non-duplicate values.
  • The returned result does not guarantee order (but it usually follows the order of first appearance).

Example 2: Extracting Unique Values from Lists and Arrays

pd.unique()It can not only handle Series, but also various array structures.

Example

import pandas as pd
import numpy as np

# 1. Extract unique values from a list
numbers = [1, 2, 3, 2, 1, 4, 5, 3, 2]
print(=== Original List ===)
print(numbers)

unique_numbers = pd.unique(numbers)
print("\n=== Unique Values from the List ===")
print(unique_numbers)
print(f"Type: {type(unique_numbers)}")

# 2. Extract unique values from a NumPy array
arr = np.array(['a', 'b', 'a', 'c', 'b', 'd'])
print("\n=== NumPy Array ===")
print(arr)
print("Unique values:", pd.unique(arr))

# 3. Extract from a DataFrame column
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Charlie', 'Alice', 'Diana'],
    'city': ['Beijing', 'Shanghai', 'Beijing', 'Beijing', 'Guangzhou']
})
print("\n=== DataFrame ===")
print(df)
print("\nUnique values of the name column:, pd.unique(df['name']))
print("Unique values of the city column:", pd.unique(df['city']))

Expected output:

=== 原始列表 ===
[1, 2, 3, 2, 1, 4, 5, 3, 2]

=== 列表的唯一值 ===
[1 2 3 4 5]
Type: <class 'numpy.ndarray'>

=== NumPy 数组 ===
['a' 'b' 'a' 'c' 'b' 'd']
唯一值: ['a' 'b' 'c' 'd']

=== DataFrame ===
      name     city
0    Alice   Beijing
1      Bob  Shanghai
2  Charlie   Beijing
3    Alice   Beijing
4   Diana  Guangzhong

=== name 列的唯一值: ['Alice' 'Bob' 'Charlie' 'Diana']
=== city 列的唯一值: ['Beijing' 'Shanghai' 'Guangzhou']

Code Explanation:

  • pd.unique()It can handle Python lists, NumPy arrays, and DataFrame columns.
  • The returned result is always a one-dimensional NumPy array.
  • When working with a DataFrame, you need to specify the specific column (using bracket syntax).

Example 3: Handling Numerical Data and Sorting

When dealing with numerical unique values, you can conveniently perform sorting and statistical analysis.

Example

import pandas as pd
import numpy as np

# 1. A Series containing duplicate numerical values
scores = pd.Series([85, 90, 78, 85, 92, 90, 78, 88, 95])

print(=== Original Scores ===)
print(scores)

# 2. Get unique values and sort them
unique_scores = pd.unique(scores)
print("\n=== Unique Values (Unsorted) ===)
print(unique_scores)

# 3. Sorted unique values
unique_sorted = np.unique(unique_scores)
print("\n=== Unique Values (Sorted) ===)
print(unique_sorted)

# 4. Count the occurrences of each unique value
print("\n=== Unique Values and Frequencies ===)
value_counts = pd.Series(scores).value_counts()
print(value_counts)

# 5. Get statistical information for the unique values
print("\n=== Basic Statistics ===)
print(f"Number of unique values: {len(unique_scores)}")
print(f"Minimum value: {unique_scores.min()}")
print(f"Maximum value: {unique_scores.max()}")
print(f"Average value: {unique_scores.mean():.2f}")

Expected output:

>--- 原始分数 ---
0    85
1    90
2    78
85、90、78 各出现 2 次,88、95 各出现 1 次。
=== 唯一值及频次 ===
85    2
90    2
78    2
88    1
95    1
dtype: int

=== 基本统计 ===
唯一值数量: 5
最小值: 78
最大值: 95
平均值: 86.40

Code Explanation:

  • You can usenp.unique()to sort the results.
  • Combined withvalue_counts()you can see the frequency of each unique value.
  • The unique value array can be used for numerical calculations like a normal array.

Example 4: Handling Missing Values

pd.unique()It will return missing values (NaN) as a valid unique value.

Example

import pandas as pd
import numpy as np

# 1. Data containing missing values
data = pd.Series(['a', 'b', np.nan, 'a', None, 'c', np.nan, 'b'])

print(=== Data with Missing Values ===)
print(data)

# 2. Get unique values (NaN will be treated as a unique value)
unique_with_nan = pd.unique(data)
print("\n=== Unique Values Including NaN ===)
print(unique_with_nan)
print(f"Number of unique values (including NaN): {len(unique_with_nan)}")

# 3. Unique values after excluding NaN
unique_no_nan = pd.unique(data[data.notna()])
print("\n=== Unique Values Excluding NaN ===)
print(unique_no_nan)
print(f"Number of unique values (excluding NaN): {len(unique_no_nan)}")

# 4. Use isna() to check whether NaN exists
has_nan = pd.isna(unique_with_nan).any()
print(f"\nWhether NaN exists: {has_nan}")

Expected output:

=== 包含缺失值的数据 ===
0       a
1       b
2     NaN
3       a
4    None
5       different 'c' appears twice: None and NaN are both treated as missing values.
print("\n=== 唯一值数量(含 NaN): 3 ===")
print("['a', 'b', nan]")

=== 排除 NaN 后的唯一值 ===
['a', 'b', 'c']
唯一值数量(不含 NaN): 3

Code Explanation:

  • Both NaN and None are treated as missing values and count as only one in the unique values.
  • You can usedata[data.notna()]to first filter out missing values, then extract unique values.
  • pd.isna()You can check whether the unique value array contains NaN.

Tip: pd.unique()The return value is a NumPy array. If you need to return a pandas object (such as a Series), you can useSeries.unique()method, which will return a Series.

Pandas 常用函数Common Pandas Functions

Other Extensions