Pandas Series.median() Function

Pandas 常用函数Pandas Common Functions


Series.median()Is a function in Pandas used to calculate the median (middle value) of a Series. The median is the value located in the middle after sorting the data. It is not affected by extreme values and is a robust indicator of the central tendency of data.

When there are outliers or skewed distributions in the data, the median can more accurately reflect the central position of the data than the mean. It is often used in scenarios such as income data analysis, housing price statistics, and grade analysis.


Basic Syntax and Parameters

median()Is a member function of the Series object, called directly via the dot operator.

Syntax Format

Series.median(axis=None, skipna=True, level=None, numeric_only=None, **kwargs)

Parameter Description

Parameter Type Description Default Value
axis int Specify the axis. A Series has only one row of data, so this parameter is mainly for compatibility with DataFrame. None
skipna bool If True, NaN values are skipped during calculation; if False, the result returns NaN when encountering NaN. True
level int or str If the Series has a MultiIndex, specify the level to calculate. None
numeric_only bool If True, only numeric data is calculated; otherwise, it attempts to convert to numeric. False

Return Value

  • Return Type:float
  • DescriptionReturns the median of all elements in the Series. If the number of elements is even, returns the average of the middle two elements.

Examples

Let's thoroughly master ... through a series of examples from simple to complex.Series.median()the usage.

Example 1: Basic Usage - Median of an Odd Number of Elements

For data with an odd number of elements, the median is the value located in the middle after sorting.

Example

import pandas as pd

# Create a Series containing employee incomes
# Simulate the monthly income of 7 employees in a department (unit: thousand yuan)
income = pd.Series([5, 6, 7, 8, 9, 10, 50])

# Calculate the median
median_income = income.median()

print("Employee monthly income (thousand yuan):")
print(income)
print()
print(f"Average income: {income.mean():.2f} thousand yuan")
print(f"Median income: {median_income} thousand yuan")
print()
print("Note: The average is greatly affected by the extreme value (50 thousand yuan), while the median better reflects the true level.")

Output:

员工月收入(千元):
0     5
1     6
2     7
3     8
4     9
5    10
6    50
dtype: int64

平均收入:13.57 千元
中位数收入:8.0 千元

Code Analysis:

  • After sorting the data: [5, 6, 7, 8, 9, 10, 50]
  • The middle position is the 4th element (index 3), with a value of 8.
  • Because there is an extreme value (50), the average (13.57) is pulled up, while the median (8) better reflects the income level of most people.

Example 2: Median of an Even Number of Elements

For data with an even number of elements, the median is the average of the middle two elements.

Example

import pandas as pd

# Create a Series with 6 elements
# Simulate the exam scores of 6 students
scores = pd.Series([75, 82, 88, 92, 95, 100])

# Calculate the median
median_score = scores.median()

print("Student exam scores:")
print(scores)
print()
print(f"Average score: {scores.mean():.2f}")
print(f"Median score: {median_score}")

Output:

学生考试成绩:
0    75
1    82
2    88
3    92
4    95
5   100
dtype: int64

平均成绩:88.67
中位数成绩:90.0

Code Analysis:

  • After sorting the data: [75, 82, 88, 92, 95, 100]
  • The middle two elements are 88 and 92 (indexes 2 and 3).
  • Median = (88 + 92) / 2 = 90.

Example 3: Handling Data Containing Missing Values

skipnaThe parameter determines how missing values are handled.

Example

import pandas as pd
import numpy as np

# Create a Series with missing values
data_with_nan = pd.Series([10, 20, np.nan, 30, 40, np.nan, 50])

print("Data with missing values:")
print(data_with_nan)
print()

# By default skipna=True, skip NaN when calculating the median
median_skipna = data_with_nan.median()
print(f"Median when skipna=True (default): {median_skipna}")

# Set skipna=False
median_no_skipna = data_with_nan.median(skipna=False)
print(f"Median when skipna=False: {median_no_skipna}")

Output:

包含缺失值的数据:
0    10.0
1    20.0
2       NaN
3    30.0
4    40.0
5       NaN
6    50.0
dtype: float64

skipna=True(默认)时的中位数:30.0
skipna=False 时的中位数:nan

Code Analysis:

  • Valid data: [10, 20, 30, 40, 50], 5 elements in total.
  • After sorting, take the middle value: 30.
  • skipna=FalseWhen ..., as long as there is a NaN, it returns NaN.

Example 4: Application Scenarios Comparing Mean and Median

Demonstrates the advantage of the median when dealing with skewed data.

Example

import pandas as pd

# Simulate the monthly income data of 10 households in a residential community
# Most people have incomes between 3000-5000 yuan, but there are a few high-income individuals
income_data = pd.Series([3000, 3500, 3800, 4000, 4200, 4500, 4800, 5000, 8000, 50000])

print("Monthly income data of community residents (yuan):")
print(income_data)
print()

mean_income = income_data.mean()
median_income = income_data.median()

print(f"Average: {mean_income:.2f} yuan")
print(f"Median: {median_income} yuan")
print()

print("Analysis:")
print("The average (10380 yuan) is greatly pulled up by the extreme value (50000 yuan).")
print("The median (4350 yuan) better reflects the true income level of most households.")
print("This is why statistical departments more commonly use the median when publishing income data.")

Output:

小区住户月收入数据(元):
0     3000
1     3500
2     3800
3     4000
4     4200
5     4500
6     4800
7     5000
8     8000
9    50000
dtype: int64

平均值:10380.00 元
中位数:4350.0 元

Notes

  • The median is insensitive to extreme values and is more robust than the mean.
  • When the number of data points is even, the median is the average of the middle two numbers.
  • When the data has a significantly skewed distribution, the median better reflects the central position of the data than the mean.
  • skipnaThe behavior of the parameter ismean()the same.

Summary

Series.median()It is an important function for describing the central tendency of data. Its main features include:

  • Not affected by extreme values, the calculation result is robust.
  • For an even number of elements, returns the average of the middle two elements.
  • Especially useful in data analysis with skewed distributions such as income and housing prices.
  • The syntax and parameters aremean()exactly the same, easy to learn and use.

In practical data analysis, it is recommended to calculate both the mean and the median to more comprehensively understand the distribution characteristics of the data. When the difference between the two is large, it indicates that the data may have a skewed distribution or outliers.

Pandas 常用函数Pandas Common Functions

Other Extensions