Pandas Series.mean() Function

Pandas 常用函数Common Pandas Functions


Series.mean()is a function in Pandas used to calculate the mean of all elements in a Series. The mean is one of the most commonly used statistical indicators, representing the central tendency of a set of data.

Whether calculating average salary, average sales, or average grades,mean()it can help you complete the task quickly. It is a core indicator for describing the central tendency of a dataset.


Basic Syntax and Parameters

mean()is a member function of the Series object, called directly via the dot operator.

Syntax Format

Series.mean(axis=None, skipna=True, level=None, numeric_only=None, **kwargs)

Parameter Description

Parameter Type Description Default Value
axis int Specifies the axis. A Series has only one row of data, so this parameter is mainly for compatibility with DataFrame. None
skipna bool If True, NaN values are skipped during calculation; if False, the result returns NaN when encountering NaN. True
level int or str If the Series has a MultiIndex, specifies which level to calculate. None
numeric_only bool If True, only numeric data is calculated; otherwise, it attempts to convert to numeric. False

Return Value

  • Return Type:float(Even if the input is an integer, the return value is converted to a float)
  • Description: Returns the mean of all valid elements in the Series. If all elements are NaN, returns NaN.

Examples

Let us thoroughly master through a series of examples from simple to complex,Series.mean()the usage of.

Example 1: Basic Usage - Calculate the Average Score of Student Grades

The most basic usage is to create a numeric Series and then callmean()to calculate the mean.

Example

import pandas as pd

# Create a Series containing student grades
# Simulate the exam scores of 8 students in a class
scores = pd.Series([85, 92, 78, 90, 88, 76, 95, 82])

# Calculate the average grade
average_score = scores.mean()

print("Student exam scores:")
print(scores)
print()
print(f"Average grade: {average_score:.2f} points")

Output:

学生考试成绩:
0    85
1    92
2    78
3    90
4    88
5    76
6    95
7    82
dtype: int64

平均成绩:85.75 分

Code Explanation:

  • Created a Series containing the grades of 8 students.
  • mean()Calculates the sum of all grades (685) divided by the number of students (8), yielding the mean 85.75.
  • Note: The return value is a float, even if the input is an integer.

Example 2: Handling Data with Missing Values

Real data often contains missing values,skipnathe parameter determines how to handle these missing values.

Example

import pandas as pd
import numpy as np

# Create a Series containing missing values
# Simulate a student absent from some exams due to illness
scores_with_nan = pd.Series([85, 92, np.nan, 90, 88, np.nan, 95, 82])

print("Data including absent exam scores:")
print(scores_with_nan)
print()

# Default skipna=True, skip NaN when calculating the mean
average_skipna = scores_with_nan.mean()
print(f"Average grade with skipna=True (default): {average_skipna:.2f}")

# Set skipna=False, return NaN when encountering NaN
average_no_skipna = scores_with_nan.mean(skipna=False)
print(f"Average grade with skipna=False: {average_no_skipna}")

Output:

包含缺考成绩的数据:
0    85.0
1    92.0
2       NaN
3    90.0
4    88.0
5       NaN
6    95.0
7    82.0
dtype: float64

skipna=True(默认)时的平均成绩:88.67
skipna=False 时的平均成绩:nan

Code Explanation:

  • Whenskipna=TrueWhen (default value), only the mean of valid values is calculated.
  • Calculation process: (85 + 92 + 90 + 88 + 95 + 82) / 6 = 532 / 6 = 88.67
  • Whenskipna=FalseWhen , as long as NaN exists, the result returns NaN.

Example 3: Using a MultiIndex Series

For a Series with a hierarchical index (MultiIndex), you can specify which level's mean to calculate.

Example

import pandas as pd

# Create a Series with a MultiIndex
# Simulate grades of students from two classes
multi_index_scores = pd.Series(
    [85, 92, 78, 90, 88, 76, 95, 82],
    index=pd.MultiIndex.from_tuples([
        ('Class A', 'Student 1'), ('Class A', 'Student 2'), ('Class A', 'Student 3'), ('Class A', 'Student 4'),
        ('Class B', 'Student 1'), ('Class B', 'Student 2'), ('Class B', 'Student 3'), ('Class B', 'Student 4')
    ],
    names=['Class', 'Student']
    )
)

print("Student grades from two classes:")
print(multi_index_scores)
print()

# Calculate the average grade of all students
overall_mean = multi_index_scores.mean()
print(f"Average grade of all students: {overall_mean:.2f}")
print()

# Calculate the average grade by class (using the level parameter)
class_a_mean = multi_index_scores.xs('Class A', level='Class').mean()
class_b_mean = multi_index_scores.xs('Class B', level='Class').mean()

print(f"Average grade of Class A students: {class_a_mean:.2f}")
print(f"Average grade of Class B students: {class_b_mean:.2f}")

Output:

两个班级的学生成绩:
班级   学生
A班    学生1    85
       学生2    92
       学生3    78
       学生4    90
B班    学生1    88
       学生2    76
       学生3    95
       学生4    82
dtype: int64

所有学生平均成绩:85.75

A班学生平均成绩:86.25
B班学生平均成绩:85.25

Code Explanation:

  • Created a Series with two index levels, named "Class" and "Student" respectively.
  • Directly callingmean()will calculate the mean of all elements.
  • Using thexs()method, you can filter data by a certain level and then calculate the mean.

Example 4: Application in Real Data Analysis

Combined with real business scenarios, demonstratingmean()typical applications of.

Example

import pandas as pd
import numpy as np

# Create simulated daily closing price data for stocks
stock_prices = pd.Series([
    100.5, 102.3, 98.7, 101.2, 103.5,
    105.2, 103.8, 107.1, 108.3, 106.9
], index=['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday',
          'Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday'])

# Create data for two weeks
week1 = stock_prices[:5]
week2 = stock_prices[5:]

print("Week 1 stock prices:")
print(week1)
print(f"Week 1 average price: {week1.mean():.2f}")
print()

print("Week 2 stock prices:")
print(week2)
print(f"Week 2 average price: {week2.mean():.2f}")
print()

# Calculate the difference in average prices between the two weeks
price_diff = week2.mean() - week1.mean()
print(f"Average price difference between two weeks: {price_diff:.2f}")

Output:

第一周股票价格:
周一    100.5
周二    102.3
周三     98.7
周四    101.2
周五    103.5
dtype: float64
第一周平均价格:101.24

第二周股票价格:
周一    105.2
周二    103.8
周三    107.1
周四    108.3
周五    106.9
dtype: float64
第二周平均价格:106.26

两周平均价格差异:5.02

Notes

  • mean()The return value is always a float, even if all input data are integers.
  • The mean is sensitive to extreme values (outliers); extreme values can significantly affect the calculation result.
  • If there are many missing values in the data, usingskipna=Trueyou should be aware that the mean is calculated based on only part of the data.
  • For Series containing non-numeric data, you need to first clean the data or use thenumeric_only=Trueparameter.

Summary

Series.mean()is one of the most commonly used statistical functions in Pandas. Its main features include:

  • Simple and easy to use, called directly via the dot operator.
  • Supports handling missing values through theskipnaparameter.
  • Supports MultiIndex, can calculate by selecting a level.
  • The return value is always a float, ensuring precision.

The mean is an important indicator for describing the central tendency of a dataset, but when extreme values exist in the data, you may need to consider using the median (median()) to better reflect the central position of the data.

Pandas 常用函数Common Pandas Functions

Other Extensions