Pandas Series.quantile() Function
Series.quantile()It is a function in Pandas used to calculate quantiles of a Series. Quantiles are values that divide ordered data into several equal parts; common ones include quartiles (25%, 50%, 75%), the median (50%), and so on.
Quantiles are important indicators for describing data distribution. They help understand the distribution shape and identify outliers. They are widely used in statistical analysis, score ranking, income analysis, and other scenarios.
Basic Syntax and Parameters
quantile()It is a member function of the Series object, called directly using the dot operator.
Syntax Format
Series.quantile(q=0.5, interpolation='linear', numeric_only=True, closed='both')
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| q | float or array-like | The quantile value(s), ranging from 0 to 1. Can be a single value or a list of multiple values. | 0.5 |
| interpolation | str | The interpolation method used when the quantile lies between two values. Optional values: 'linear', 'lower', 'higher', 'nearest', 'midpoint'. | 'linear' |
| numeric_only | bool | If True, only numeric data is calculated. | True |
| closed | str | Used in DataFrame to determine the closure of the interval. Not commonly used in Series. | 'both' |
Return Value
- Return Type:
floatorSeries - DescriptionReturns the value at the specified quantile. If q is a single value, returns a float; if q is a list, returns a Series.
Examples
Let's thoroughly master, through a series of examples from simple to complex,Series.quantile()the usage of Series.quantile().
Example 1: Basic Usage - Calculating the Median
The median is the 50% quantile and is the most commonly used quantile.
Example
# Create a Series containing student scores
scores = pd.Series([65, 70, 72, 75, 78, 80, 82, 85, 88, 90, 92, 95])
print("Student scores:")
print(scores)
print()
# Calculate the median (50% quantile)
median_score = scores.quantile(0.5)
print(f"Median: {median_score}")
print(f"Using median() function: {scores.median()}")
print()
print("Analysis: 50% of student scores are less than or equal to 80.")
Output:
学生成绩: 0 65 1 70 2 72 3 75 4 75 5 80 6 80 7 85 8 88 9 90 10 92 11 95 dtype: int64 中位数:80.0 中位数(使用 median() 函数):80.0
Code explanation:
quantile(0.5)Equivalent tomedian()。- The median divides the data into two halves; 50% of the data is less than or equal to the median.
Example 2: Calculating Quartiles
Quartiles divide the data into four equal parts: 25%, 50%, 75%.
Example
# Create a Series containing employee income
income = pd.Series([3000, 3500, 3800, 4000, 4200, 4500, 5000, 5500, 6000, 8000, 15000])
print("Employee monthly income data (yuan):")
print(income)
print()
# Calculate the three quartiles
q1 = income.quantile(0.25) # First quartile (25%)
q2 = income.quantile(0.50) # Second quartile / median (50%)
q3 = income.quantile(0.75) # Third quartile (75%)
print(f"First quartile Q1 (25%): {q1} yuan")
print(f"Second quartile Q2 (50%): {q2} yuan")
print(f"Third quartile Q3 (75%): {q3} yuan")
print()
# Calculate the interquartile range (IQR)
iqr = q3 - q1
print(f"Interquartile range IQR: {iqr} yuan")
print()
print("Analysis:")
print("- 25% of employees earn less than or equal to 4000 yuan")
print("- 50% of employees earn less than or equal to 5000 yuan")
print("- 75% of employees earn less than or equal to 6000 yuan")
print("- The larger the IQR, the more dispersed the data")
Output:
员工月收入数据(元): 1 3000 2 低于 3800 3 4000 4 4200 5 4500 6 5000 7 5500 8 6000 9 8000 10 15000 dtype: int64 Q1(25% 分位):4000.0 元 Q2(50% 分位):5000.0 元 Q3(75% 分位):6000.0 元 四分位距 IQR = 6000 - 4000 = 2000 元
Example 3: Calculating Multiple Quantiles at Once
You can calculate multiple quantiles at once.
Example
# Create data
data = pd.Series([10, 20, 30, 40, 50, 60, 70, 80, 90, 100])
print("Data:")
print(data)
print()
# Calculate multiple quantiles
percentiles = data.quantile([0, 0.1, 0.25, 0.5, 0.75, 0.9, 1.0])
print("Quantiles:")
print(percentiles)
print()
# The quantile can also be specified as a percentage string (single value only)
print(f"Using string '50%': {data.quantile('50%')}")
print(f"Using string '0.5': {data.quantile(0.5)}")
Output:
数据: 0 10 1 20 2 30 3 40 4 50 量值:60 5 70 6 80 7 90 8 100 dtype: int 参数 多个分位数: 0.00 10.0 0.10 19.0 0.25 32.5 0.50 50.0 0.75 67.5 0.90 81.0 1.00 100.0 dtype: float64
Example 4: The Role of the interpolation Parameter
When the quantile lies between two values, different interpolation methods produce different results.
Example
# Create a Series with 6 elements
data = pd.Series([10, 20, 30, 40, 50, 60])
print("Data:", data.values)
print()
# Calculate the 30% quantile (between 20 and 30)
# position = (n-1) * q = 5 * 0.3 = 1.5
print("30% quantile using different interpolation methods:")
linear = data.quantile(0.3, interpolation='linear')
print(f"linear (linear interpolation, default): {linear}")
lower = data.quantile(0.3, interpolation='lower')
print(f"lower (take the smaller value): {lower}")
higher = data.quantile(0.3, interpolation='higher')
print(f"higher (take the larger value): {higher}")
nearest = data.quantile(0.3, interpolation='nearest')
print(f"nearest (take the nearest value): {nearest}")
midpoint = data.quantile(0.3, interpolation='midpoint')
print(f"midpoint (take the midpoint): {midpoint}")
Output:
数据:[10, 20, 30, 40, 50, 60] 位置计算:(n-1) * q = 5 * 0.3 = 1.5,表示在 20 和 30 之间 不同插值方法的结果: - linear(线性插值,默认):23.0 - lower(向下取整):20.0 - higher(向上取整):30.0 - nearest(最近邻):20.0 - midpoint(中间值):25.0
Code explanation:
interpolation='linear'Performs linear interpolation between the two values, 20 + (30-20)*0.5 = 23.interpolation='lower'Takes the value with the smaller index, i.e., 20.interpolation='higher'Takes the value with the larger index, i.e., 30.interpolation='nearest'Takes the value closest to the quantile position.interpolation='midpoint'Takes the midpoint of the two values, i.e., (20+30)/2 = 25.
Example 5: Using Quantiles to Identify Outliers
The interquartile range (IQR) is often used to identify outliers in data.
Example
# Create data containing some outliers
data = pd.Series([12, 15, 18, 20, 22, 25, 28, 30, 32, 150])
print("Data (with outliers):")
print(data)
print()
# Calculate quartiles
q1 = data.quantile(0.25)
q3 = data.quantile(0.75)
iqr = q3 - q1
# Calculate outlier boundaries
lower_bound = q1 - 1.5 * iqr
upper_bound = q3 + 1.5 * iqr
print(f"Q1(25%):{q1}")
print(f"Q3(75%):{q3}")
print(f"IQR:{iqr}")
print()
print(f"Lower bound of normal values: {lower_bound}")
print(f"Upper bound of normal values: {upper_bound}")
print()
# Identify outliers
outliers = data[(data upper_bound)]
print(f"Outliers: {outliers.values}")
print()
print("Analysis: 150 is an obvious outlier, exceeding the upper bound.")
Output:
数据(含异常值): 3 150 4 22 5 25 6 28 7 30 8 32 9 150 dtype: int64 Q1(25%):18.75 Q3(75%):30.25 IQR = 11.5 正常值下界:1.5 正常值上界:47.5 异常值:150 分析:150 明显超出了正常范围,是一个异常值。
Notes
- The quantile value ranges from 0 to 1.
- By default, the linear interpolation method is used to calculate quantiles.
- When q is a list, the return value is a Series rather than a single value.
- Quantiles can be used to identify outliers in data; a common method is the IQR rule (1.5 times the interquartile range).
- For large datasets, quantile calculation is very efficient.
Summary
Series.quantile()It is an important function for analyzing data distribution. Its main features include:
- Supports calculating any quantile (between 0 and 1).
- Can calculate multiple quantiles at once.
- Provides multiple interpolation methods to meet different needs.
- Very useful for identifying outliers.
In practical data analysis, quantiles are often used to understand the data distribution shape, compare different datasets, identify outliers, and so on. When combined with box plots, they can display the data distribution more intuitively.
Other Extensions
Common Pandas Functions