Pandas Series.sum() Function

Pandas 常用函数Common Pandas Functions


Series.sum()It is a function in Pandas used to calculate the sum of all elements in a Series. It is one of the most commonly used statistical functions in data analysis and can quickly obtain the sum of numeric data.

Whether calculating total sales, total profit, or total quantity,sum()it can help you complete it quickly. It is especially suitable for handling financial data, statistical reports, sales data analysis, and similar scenarios.


Basic Syntax and Parameters

sum()It is a member function of the Series object and is called directly via the dot operator.

Syntax Format

Series.sum(axis=None, skipna=True, level=None, numeric_only=None, min_count=0, **kwargs)

Parameter Description

Parameter Type Description Default Value
axis int Specifies the axis. A Series has only one row of data; this parameter is mainly for compatibility with DataFrame. None
skipna bool If True, NaN values are skipped during calculation; if False, then the result returns NaN when NaN is encountered. True
level int or str If the Series has a MultiIndex, specify the level to calculate. None
numeric_only bool If True, only numeric data is calculated; otherwise, it attempts to convert to numeric values. False
min_count int The minimum number of valid values required for calculation. If the number of valid values is less than this count, NaN is returned. 0

Return Value

  • Return Type: numeric (int、floatornumpy.nan)
  • Description: Returns the sum of all elements in the Series. If all elements are NaN (and skipna=True), returns 0.

Examples

Let us, through a series of examples from simple to complex, thoroughly masterSeries.sum()the usage of it.

Example 1: Basic Usage - Calculate the Sum of a Numeric List

The most basic usage is to create a numeric Series and then callsum()to calculate the sum.

Example

import pandas as pd

# Create a Series containing sales amounts
# Simulate daily sales data for a week
sales_data = pd.Series([1200, 1500, 1800, 900, 2100, 1600, 1350])

# Calculate total sales
total_sales = sales_data.sum()

print("Daily sales:")
print(sales_data)
print()
print(f"Total sales: {total_sales}")

Output:

每日销售额:
0    1200
1    1500
2    1800
3     900
4    2100
5    1600
6    1350
dtype: int64

总销售额:10450

Code explanation:

  • Created a Series containing 7 days of sales data, with an integer data type.
  • sum()It simply iterates over all elements and adds them up to get the total 10450.

Example 2: Handling Data Containing Missing Values

Missing values often occur in real data,skipnaand the parameter determines how to handle these missing values.

Example

import pandas as pd
import numpy as np

# Create a Series containing missing values
# Simulate missing sales data for certain dates
sales_with_nan = pd.Series([1200, 1500, np.nan, 1800, np.nan, 2100, 1600])

print("Sales data with missing values:")
print(sales_with_nan)
print()

# Default skipna=True, skip NaN and calculate the sum
total_skipna = sales_with_nan.sum()
print(f"Sum with skipna=True (default): {total_skipna}")

# Set skipna=False, return NaN when NaN is encountered
total_no_skipna = sales_with_nan.sum(skipna=False)
print(f"Sum with skipna=False: {total_no_skipna}")

Output:

包含缺失值的销售数据:
0    1200.0
1    1500.0
2       NaN
3    1800.0
4       NaN
5    2100.0
6    1600.0
dtype: float64

skipna=True(默认)时的总和:8200.0
skipna=False 时的总和:nan

Code explanation:

  • Whenskipna=TrueWhen (the default value),sum()it automatically skips NaN values and calculates the sum of only valid values.
  • Calculation process: 1200 + 1500 + 1800 + 2100 + 1600 = 8200
  • Whenskipna=FalseWhen skipna=False, as long as NaN exists, the result returns NaN. This is useful in scenarios where you need to clearly know data integrity.

Example 3: Using the min_count Parameter

min_countThe min_count parameter can set the minimum number of valid values required for calculation. This is useful when ensuring data sufficiency.

Example

import pandas as pd
import numpy as np

# Create a Series where most values are missing
sparse_data = pd.Series([np.nan, np.nan, 100, np.nan, np.nan])

print("Data with most values missing:")
print(sparse_data)
print()

# Default min_count=0, calculate as long as there are at least 0 valid values
result_default = sparse_data.sum()
print(f"Result with min_count=0 (default): {result_default}")

# Set min_count=3, require at least 3 valid values to calculate
result_min_count = sparse_data.sum(min_count=3)
print(f"Result with min_count=3: {result_min_count}")

# Set min_count=2, only need 2 valid values
result_min_count_2 = sparse_data.sum(min_count=2)
print(f"Result with min_count=2: {result_min_count_2}")

Output:

大部分缺失的数据:
0     NaN
1     NaN
2    100.0
3     NaN
4     NaN
dtype: float64

min_count=0(默认)时的结果:100.0
min_count=3 时的结果:nan
min_count=2 时的结果:100.0

Code explanation:

  • This Series has only 1 valid value (100).
  • min_count=3This means at least 3 valid values are required for calculation, but there is actually only 1, so NaN is returned.
  • min_count=2This means at least 2 valid values are required, but there is actually 1 valid value, which does not meet the condition, so NaN is returned.
  • This parameter is very useful in analysis scenarios with high data quality requirements, as it can help identify situations where valid data is insufficient.

Example 4: Application in Real Data Analysis

Combining real business scenarios, demonstratesum()its typical applications.

Example

import pandas as pd

# Create a simulated sales data DataFrame
sales_data = pd.DataFrame({
    'Month': ['Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun'],
    'East China Region': [12000, 15000, 18000, 16000, 14000, 17000],
    'North China Region': [10000, 11000, 13000, 12500, 11500, 14000],
    'South China Region': [15000, 17000, 19000, 18000, 16000, 20000]
})

print("Sales data by region for the first half of the year:")
print(sales_data)
print()

# Calculate the half-year total sales for each region
for region in ['East China Region', 'North China Region', 'South China Region']:
    total = sales_data[region].sum()
    print(f"{region} half-year total sales: {total} yuan")

print()

# Calculate the total sales for all regions
total_all = sales_data[['East China Region', 'North China Region', 'South China Region']].sum().sum()
print(f"Total half-year sales for all regions: {total_all} yuan")

Output:

上半年各区域销售数据:
   月份   华东区   华北区   华南区
0  1月  12000  10000  15000
1  2月  15000  11000  17000
2  3月  18000  13000  19000
3  4月  16000  12500  18000
4  5月  14000  11500  16000
5  6月  17000  14000  20000

华东区半年总销售额:92000 元
华北区半年总销售额:72000 元
华南区半年总销售额:105000 元

所有区域半年总销售额:269000 元

Notes

  • sum()By default, NaN values are skipped, which is the expected behavior in most cases.
  • If you need to sum a Series containing non-numeric data, you should first clean the data or use thenumeric_only=Truenumeric_only parameter.
  • min_countThe min_count parameter is very useful in data validation and quality checks, as it can help identify situations where valid data is insufficient.
  • For large datasets,sum()sum() usually performs well because it uses NumPy's vectorized operations under the hood.

Summary

Series.sum()It is one of the most basic and commonly used statistical functions in Pandas. Its main features include:

  • Simple and easy to use, called directly through the dot operator.
  • Supports handling missing values via theskipnaskipna parameter.
  • You can use themin_countmin_count parameter to set the minimum requirement for valid values.
  • It uses NumPy's optimized implementation at the underlying level, resulting in high computational efficiency.

In real data analysis,sum()it is usually used together with other aggregation functions (such asmean()、count()mean(), etc.) to conduct comprehensive statistical analysis of data.

Pandas 常用函数Common Pandas Functions

Other Extensions