Pandas pd.cut() Function
pd.cut()is used in the Pandas library tobin continuous data (Binning)function. It divides numerical data into discrete categories according to specified intervals, and is very suitable for grouped statistics and visualization in data analysis.
Binning operations are very common in data analysis, such as dividing ages into "teenager", "middle-aged", "elderly", or dividing scores into "fail", "pass", "good", "excellent", etc.
Word Meaning: cutIt means "cut". Here it refers to "cutting" a continuous data interval into multiple discrete intervals.
Basic Syntax and Parameters
pd.cut()is a top-level function of the Pandas library, used to discretize continuous variables.
Syntax Format
pd.cut(x, bins, right=True, labels=None, retbins=False, precision=3, include_lowest=False)
Parameter Description
- Parameter:
x- Type: Array-like object, such as list, Series, array, etc.
- Description: The continuous numerical data to be binned.
- Parameter:
bins- Type: Integer, scalar (interval boundaries), or IntervalIndex.
- Description: The binning method. It can be an integer (indicating how many equal-width bins to divide into) or a list (specifying specific interval boundaries).
- Parameter:
right- Type: Boolean value.
- Description: Whether to use right-closed intervals. Default is
True(i.e., in the form [a, b]). If set toFalse, then it is a left-closed right-open interval (a, b].
- Parameter:
labels- Type: Array or None.
- Description: Specify custom labels for each interval. If not specified, interval representations such as "(0, 10]" are used.
- Parameter:
retbins- Type: Boolean value.
- Description: If
True, the return result will contain the bin boundaries. Default isFalse。
- Parameter:
include_lowest- Type: Boolean value.
- Description: If
True, the first interval contains the left boundary. Default isFalse。
Function Description
- Return Value: Returns a Series of Categorical type, where each value corresponds to the interval to which it belongs.
- Effect: Maps continuous numerical values to discrete categories (intervals).
Example
Let's thoroughly master, through a series of examples from simple to complex,pd.cut()the usage.
Example 1: Basic Usage — Equal-Width Binning
Example
import numpy as np
# 1. Create numerical data
scores = pd.Series([55, 70, 82, 45, 90, 65, 78, 88, 92, 58])
print("=== Original scores ===")
print(scores)
# 2. Divide the scores into 3 equal-width intervals
result = pd.cut(scores, bins=3)
print("\n=== pd.cut(scores, bins=3) equal-width binning ===)
print(result)
# 3. View the bin boundaries
result_bins = pd.cut(scores, bins=3, retbins=True)
print("\n=== Return bin boundaries ===)
print(f"Bin boundaries: {result_bins)
Expected output:
=== 原始分数 === 0 55 1 70 2 82 3 45 4 90 5 65 6 78 7 88 8 92 9 58 === pd.cut(scores, bins=3) 等宽分箱 === 0 (44.667, 60.667] 1 (60.667, 76.333] 2 (76.333, 92.0] 3 (44.667, 60.667] 4 (76.333, 92.0] 5 (60.667, 76.333] 6 (60.667, 76.333] 7 (76.333, 92.0] 8 (76.333, 92.0] 9 (44.667, 60.667] dtype: category Categories (3, interval[float64]): [(44.667, 60.667] < (60.667, 76.333] < (76.333, 92.0] === 返回分箱边界 === 分箱边界: [44.667 60.667 76.333 92. ]
Code explanation:
bins=3Indicates that the data is automatically divided into 3 equal-width intervals.- Pandas automatically calculates the interval boundaries based on the minimum and maximum values of the data.
- The return value is of Categorical type, and each value is an Interval object.
- By default, right-closed intervals are used, i.e., in the form (a, b].
Example 2: Custom Interval Boundaries
In practical applications, we often need to customize bin boundaries based on business requirements.
Example
import numpy as np
# Create score data
scores = pd.Series([55, 70, 82, 45, 90, 65, 78, 88, 92, 58, 73, 67, 81, 49, 95])
# Custom bin boundaries: Fail, Pass, Good, Excellent
bins = [0, 60, 75, 90, 100]
labels = ['Fail', 'Pass', 'Good', 'Excellent']
result = pd.cut(scores, bins=bins, labels=labels, include_lowest=True)
print("=== Original scores ===")
print(scores.values)
print("\n=== Binning result ===)
print(result)
# Count the number of people in each grade
print("\n=== Score distribution statistics ===)
print(result.value_counts().sort_index())
Expected output:
=== 原始成绩 [55 70 82 45 90 65 78 88 92 58 73 67 81 49 95] === 分箱结果 0 及格 1 良好 2 良好 3 不及格 4 优秀 ... 12 良好 13 不及格 14 优秀 dtype: category Categories (4, interval[int64]): [不及格 < 及格 < 良好 < 优秀] === 成绩分布统计 不及格 3 及格 4 良好 5 优秀 3 dtype: int64
Code explanation:
bins=[0, 60, 75, 90, 100]Defined 4 intervals: [0-60], (60-75], (75-90], (90-100]labelsThe labels parameter specifies intuitive Chinese labels for each intervalinclude_lowest=TrueEnsure that the first interval contains the left boundary 0value_counts()Makes it easy to count the number of people in each grade
Example 3: Using Right-Closed/Left-Closed Intervals
Example
import numpy as np
# Test data
data = pd.Series([10, 20, 30, 40, 50])
# Default right-closed interval [a, b]
result1 = pd.cut(data, bins=4)
print("=== Default right-closed interval [a, b] ===")
print(result1)
# Left-closed right-open interval (a, b]
result2 = pd.cut(data, bins=4, right=False)
print("\n=== Left-closed right-open interval (a, b] ===)
print(result2)
Expected output:
=== 默认右闭区间 [a, b] 0 (7.5, 17.5] 1 (17.5, 27.5] 2 (27.5, 37.5] 3 (37.5, 47.5] 4 (47.5, 57.5] dtype: category === 左闭右开区间 (a, b] 0 [10, 17.5) 1 [17.5, 27.5) 2 [27.5, 37.5) 3 [37.5, 47.5) 4 [47.5, 57.5) dtype: category
Code explanation:
- By default, the right-closed interval is used
[a, b], i.e., includes the right boundary - Set
right=Falseto change to a left-closed right-open interval[a, b) - The open/closed mode of the interval affects the classification of boundary values.
Notes
Important notes:
- If there are values in the data that exceed
binsthe specified range, NaN will be producedNaNbinsboundaries must be monotonically increasing- When using custom labels, the number of labels must be one less than the number of boundaries.
- The binned result is of Categorical type and can be processed using string methods.
Other Extensions
Pandas Common Functions