Pandas Series.cumsum() Function
Series.cumsum()is a Pandas function used to calculate the cumulative sum of a Series. The cumulative sum means that starting from the first element, the value at each position equals the sum of all elements before that position (including the current element).
Cumulative sums are often used in financial data analysis (such as calculating cumulative returns), sales analysis (such as calculating cumulative sales), time series analysis, and other scenarios.
Basic Syntax and Parameters
cumsum()is a member function of the Series object, called directly via the dot operator.
Syntax
Series.cumsum(axis=None, skipna=True, dtype=None, out=None, **kwargs)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| axis | int | Specifies the axis; a Series has only one column, so this parameter is mainly for compatibility with DataFrame. | None |
| skipna | bool | If True, NaN values are skipped during calculation; if False, encountering NaN will break the cumulative sum. | True |
| dtype | dtype | Specifies the output data type. | None |
| out | ndarray | The array used to store the result; usually does not need to be set. | None |
Return Value
- Return Type:
Series - Description: Returns a new Series, where each position's value is the cumulative sum.
Examples
Through a series of examples from simple to complex, let us thoroughly masterSeries.cumsum()the usage.
Example 1: Basic Usage - Calculating Cumulative Sales
The cumulative sum is the total of all values up to the current position.
Example
# Create a Series containing daily sales
daily_sales = pd.Series([1000, 1500, 1200, 1800, 2000, 900, 1600])
print("Daily sales (yuan):")
print(daily_sales)
print()
# Calculate cumulative sales
cumulative_sales = daily_sales.cumsum()
print("Cumulative sales (yuan):")
print(cumulative_sales)
print()
print("Analysis:")
print("- Day 1 cumulative: 1000")
print("- Day 2 cumulative: 1000 + 1500 = 2500")
print("- Day 3 cumulative: 1000 + 1500 + 1200 = 3700")
print("And so on...")
Output:
每日销售额(元): 0 1000 1 1500 2 1200 3 1800 4 2000 5 900 6 1600 dtype: int64 累计销售额(元): 0 1000 1 2500 2 3700 3 5500 4 7500 5 8400 6 10000 dtype: int64 分析: - 第1天的累计销售就是当天的销售额 - 第2天的累计销售额 = 第1天 + 第2天的销售额 - 以此类推...
Code explanation:
- Cumulative sum at position 1 = 1000
- Cumulative sum at position 2 = 1000 + 1500 = 2500
- Cumulative sum at position 3 = 1000 + 1500 + 1200 = 3700
- And so on...
Example 2: Handling Data with Missing Values
skipnaThe parameter determines how missing values are handled.
Example
import numpy as np
# Create a Series containing missing values
data_with_nan = pd.Series([10, 20, np.nan, 30, 40])
print("Original data (with missing values):")
print(data_with_nan)
print()
# Default skipna=True, skip NaN when calculating cumulative sum
cumulative_skipna = data_with_nan.cumsum()
print("Cumulative sum with skipna=True (default):")
print(cumulative_skipna)
print()
# Set skipna=False, NaN positions will break the cumulative sum
cumulative_no_skipna = data_with_nan.cumsum(skipna=False)
print("Cumulative sum with skipna=False:")
print(cumulative_no_skipna)
Output:
原始数据(含缺失值): 0 10.0 1 20.0 2 NaN 3 30.0 4 40.0 dtype: float64 skipna=True(默认)的累计和: 0 10.0 1 30.0 2 NaN 3 60.0 4 100.0 dtype: float64 skipna=False 的累计和: 0 10.0 1 30.0 2 NaN 3 NaN 4 DataFrame
Code explanation:
- When
skipna=TrueWhen set to True, NaN positions are skipped, and subsequent cumulative sums continue based on valid values. - Calculation: 10 → 30 → (skip) → 60 → 100
- When
skipna=FalseWhen set to False, all positions after encountering NaN become NaN.
Example 3: Analyzing Together with Original Data
In practical analysis, it is often necessary to view both the original data and the cumulative data at the same time.
Example
# Create stock price data
stock_prices = pd.Series([100, 102, 98, 105, 103, 108, 110])
# Calculate daily changes
daily_change = stock_prices.diff()
# Calculate cumulative changes
cumulative_change = stock_prices.cumsum() - stock_prices.iloc[0] * len(stock_prices) + stock_prices
# Or more simply, use cumsum to calculate the cumulative change relative to the initial price
print("Stock price data:")
print(stock_prices.values)
print()
# Calculate the price change for the day (compared with the previous day)
print("Daily changes:")
print(daily_change.fillna(0).values)
print()
# Another way: calculate cumulative returns
cumulative_return = ((stock_prices / stock_prices.iloc[0]) - 1) * 100
print("Cumulative returns (%):")
print(cumulative_return.values)
print()
# Cumulative value starting from 100 yuan
initial_price = 100
cumulative_value = initial_price + stock_prices.cumsum() - stock_prices.iloc[0]
print("Cumulative value (assuming initial 100 yuan):")
print(cumulative_value.values)
Output:
股价数据:[100, 102, 98, 105, 103, 108, 110] 每日涨跌(与前一天相比): [ 0., 2., -4., 7., -2., 5., 2.] 累计收益(%): [0.0, 2.0, -2.0, 5.0, 3.0, 8.0, 10.0] 累计价值(假设初始 100 元): 0 100 1 202 2 300 3 405 4 508 5 616 6 726 dtype: int64 分析:第 7 天时,股价从 100 涨到 110,累计收益率为 10%。
Example 4: Practical Application - Monthly Cumulative Data
Demonstrates the typical application of cumsum in financial analysis.
Example
# Create monthly sales data
monthly_sales = pd.Series({
'January': 50000,
'February': 62000,
'March': 58000,
'April': 70000,
'May': 75000,
'June': 80000
})
print("Monthly sales (yuan):")
print(monthly_sales)
print()
# Calculate monthly cumulative sales
cumulative_sales = monthly_sales.cumsum()
# Create summary table
summary = pd.DataFrame({
'Monthly sales': monthly_sales,
'Cumulative sales': cumulative_sales,
'Annual target completion ratio': (cumulative_sales / 500000 * 100).round(1)
})
print("Sales summary table:")
print(summary)
print()
print(f"First half cumulative sales: {cumulative_sales.iloc[-1]} yuan")
print(f"Annual target completion rate: {cumulative_sales.iloc[-1]/500000*100:.1f}%")
Output:
月度销售额(元):
1月 50000
2月 62000
累计 112000
4月 70000
5月 75000
6月 80000
dtype: int64
销售汇总表:
月度销售额 累计销售额 完成年度目标(%)
1月 50000 50000 10.0
2月 62000 112000 22.4
3月 58000 170000 34.0
4月 70000 240000 累计 34.0%
5月 75000 315000 63.0
6月 累计 395000 79.0
dtype: int64
分析:6 月底累计销售额为 39.5 万元,已完成年度目标的 79.0%。
Example 5: Application in DataFrame
cumsum can also be used in a DataFrame column-wise or row-wise.
Example
# Create a DataFrame containing sales of multiple products
sales_data = pd.DataFrame({
'Product A': [100, 150, 120, 180],
'Product B': [80, 90, 110, 130],
'Product C': [50, 60, 70, 80]
}, index=['January', 'February', 'March', 'April'])
print("Monthly sales data:")
print(sales_data)
print()
# Calculate cumulative sales by column (default)
cumulative_by_col = sales_data.cumsum()
print("Cumulative sales by product:")
print(cumulative_by_col)
print()
# Calculate cumulative sales by row
cumulative_by_row = sales_data.cumsum(axis=1)
print("Cumulative sales by month:")
print(cumulative_by_row)
Output:
月度销售数据:
产品A 产品B 产品C
1月 100 80 50
2月 150 90 60
3月 120
4月 180 130 80
按产品累计销量:
产品A 产品B 产品C
1月 100 80 50
2月 250 170 110
3月 370 280 180
4月 550 410 260
按月度累计销量:
产品A 产品B 产品C
1月 100.0 180.0 230.0
2月 250.0 340.0 400.0
Code explanation:
- By default,
axis=0the cumulative sum is calculated along the column direction. - Setting
axis=1calculates the cumulative sum along the row direction.
Notes
- cumsum returns a new Series and does not modify the original data.
- By default, skipna=True, and NaN values are skipped.
- When skipna=False, encountering NaN will cause all subsequent values to become NaN.
- In a DataFrame, the axis parameter can be used to control the cumulative direction.
- The calculation of the cumulative sum is order-dependent and cannot be parallelized.
Summary
Series.cumsum()is a fundamental function in time series analysis. Its main features include:
- Calculates the cumulative value from the beginning to the current position.
- Supports handling missing values, controlled via the skipna parameter.
- Supports row-wise or column-wise calculation in a DataFrame.
- Widely used in financial analysis, sales statistics, progress tracking, and other scenarios.
The cumulative sum helps us understand the trend of data over time and is an important tool in data analysis. Used together with other cumulative functions (such as cumprod, cummax, cummin), it can provide more comprehensive data insights.
Other Extensions
Common Pandas Functions