Pandas Series.std() Function
Series.std()It is a function in Pandas for calculating the standard deviation of a Series. Standard deviation is a measure of data dispersion, representing the average degree of deviation between data and its mean.
The larger the standard deviation, the more dispersed the data; the smaller the standard deviation, the more concentrated the data. It is often used in scenarios such as quality control, risk assessment, and exam score analysis.
Basic Syntax and Parameters
std()It is a member function of the Series object and is called directly via the dot operator.
Syntax Format
Series.std(axis=None, skipna=True, level=None, numeric_only=None, ddof=1, **kwargs)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| axis | int | Specifies the axis. Series has only one row of data; this parameter is mainly for compatibility with DataFrame. | None |
| skipna | bool | If True, skip NaN values during calculation; if False, the result will return NaN when encountering NaN. | True |
| level | int or str | If the Series has a MultiIndex, specify the level to calculate. | None |
| numeric_only | bool | If True, only operate on numeric data; otherwise, attempt to convert to numeric. | False |
| ddof | int | Delta degrees of freedom adjustment parameter. ddof=1 uses sample standard deviation (n-1), ddof=0 uses population standard deviation (n). | 1 |
Return Value
- Return Type:
float - Description: Returns the standard deviation of the elements in the Series. By default, uses the sample standard deviation (divided by n-1).
Examples
Let's, through a series of examples from simple to complex, thoroughly masterSeries.std()the usage.
Example 1: Basic Usage - Understanding the Concept of Standard Deviation
Standard deviation measures the dispersion of data; the larger the value, the more dispersed the data.
Example
import numpy as np
# Two sets of score data
# Group A: scores are relatively concentrated
group_a = pd.Series([85, 86, 87, 88, 89])
# Group B: scores are relatively dispersed
group_b = pd.Series([70, 75, 85, 95, 100])
print("Group A scores (more concentrated):")
print(group_a)
print(f"Mean: {group_a.mean():.2f}")
print(f"Standard deviation: {group_a.std():.2f}")
print()
print("Group B scores (more dispersed):")
print(group_b)
print(f"Mean: {group_b.mean():.2f}")
print(f"Standard deviation: {group_b.std():.2f}")
print()
print("Analysis: Although the two groups have the same mean (85), Group B has a larger standard deviation, indicating greater score differences.")
Output:
A组成绩(较集中): 0 85 1 86 2 87 3 88 4 89 dtype: int64 平均值:85.00 标准差:1.58 B组成绩(较分散): 0 70 1 75 2 85 3 95 4 100 dtype: int64 平均值:85.00 标准差:12.50 分析:虽然两组平均值相同(85),但 B 组标准差更大,说明成绩差异更大。
Code Explanation:
- Group A's standard deviation is about 1.58, very concentrated.
- Group B's standard deviation is about 12.50, much more dispersed.
- This shows that even with the same mean, the distribution of data can be completely different.
Example 2: The Role of the ddof Parameter
ddofThe parameter controls whether to use sample standard deviation or population standard deviation.
import numpy as np
# Create a dataset
data = pd.Series([2, 4, 4, 4, 5, 5, 7, 9])
print("Data:")
print(data)
print()
# Default ddof=1, uses sample standard deviation (divided by n-1)
sample_std = data.std(ddof=1)
print(f"Sample standard deviation (ddof=1): {sample_std:.4f}")
# ddof=0, uses population standard deviation (divided by n)
population_std = data.std(ddof=0)
print(f"Population standard deviation (ddof=0): {population_std:.4f}")
print()
print("Explanation:")
print("Sample standard deviation = sqrt(sum((x-mean)^2) / (n-1))")
print("Population standard deviation = sqrt(sum((x-mean)^2) / n)")
print("When the amount of data is large, the difference between the two is very small.")
Output:
数据: 0 2 1 4 2 4 3 4 4 5 5 5 6 7 7 9 dtype: int64 样本标准差(ddof=1):2.2678 总体标准差(ddof=0):2.1213
Code Explanation:
- Sample standard deviation (ddof=1) uses n-1 as the divisor and is suitable for samples drawn from a population.
- Population standard deviation (ddof=0) uses n as the divisor and is suitable for the entire dataset.
- Pandas defaults to ddof=1, i.e., sample standard deviation.
Example 3: Handling Data Containing Missing Values
Example
import numpy as np
# Create a Series containing missing values
data_with_nan = pd.Series([10, 20, np.nan, 30, 40, np.nan, 50])
print("Data containing missing values:")
print(data_with_nan)
print()
# Default skipna=True, skips NaN when calculating standard deviation
std_skipna = data_with_nan.std()
print(f"Standard deviation when skipna=True (default): {std_skipna:.4f}")
# Set skipna=False
std_no_skipna = data_with_nan.std(skipna=False)
print(f"Standard deviation when skipna=False: {std_no_skipna}")
Output:
包含缺失值的数据: 0 10.0 1 20.0 2 NaN 3 30.0 4 40.0 5 NaN 6 50.0 dtype: float64 skipna=True(默认)时的标准差:15.8114 skipna=False 时的标准差:nan
Example 4: Practical Application - Stock Return Volatility Analysis
Standard deviation is commonly used in finance to measure risk.
Example
# Simulate 10 days of daily returns (%) for two stocks
stock_a = pd.Series([1.2, 0.8, -0.5, 1.5, 0.3, -0.2, 1.0, 0.7, -0.3, 0.5])
stock_b = pd.Series([3.5, -2.0, 4.2, -1.5, 2.8, -3.0, 1.2, -0.8, 3.0, -2.4])
print("Stock A daily returns (%):")
print(stock_a)
print(f"Average return: {stock_a.mean():.2f}%")
print(f"Volatility (standard deviation): {stock_a.std():.2f}%")
print()
print("Stock B daily returns (%):")
print(stock_b)
print(f"Average return: {stock_b.mean():.2f}%")
print(f"Volatility (standard deviation): {stock_b.std():.2f}%")
print()
print("Analysis:")
print("Stock B has a larger standard deviation, indicating more volatile returns and higher risk.")
print("Although their average returns may be similar, the level of risk is different.")
Output:
股票 A 日收益率(%): 0 1.2 1 0.8 2 -0.5 3 1.5 4 0.3 5 -0.2 6 1.0 7 0.7 8 -0.3 9 0.5 dtype: int64 平均收益率:0.60% 波动率(标准差):0.68% 股票 B 日收益率(%): 0 3.5 1 -2.0 2 4.2 3 -1.5 4 2.8 5 -3.0 6 1.2 7 -0.8 8 3.0 9 -2.4 dtype: int64 平均收益率:0.60% 波动率(标准差):2.59% 分析:股票 B 的波动率是股票 A 的约 4 倍,说明风险高得多。
Notes
- By default, the sample standard deviation (ddof=1) is used, suitable for statistical analysis.
- If you need to calculate the population standard deviation, set ddof=0.
- The unit of standard deviation is the same as the original data, making it easy to understand intuitively.
- The larger the standard deviation, the more dispersed the data; the smaller the standard deviation, the more concentrated the data.
Summary
Series.std()It is an important function for measuring data dispersion. Its main features include:
- By default, it uses the sample standard deviation (ddof=1), which is more suitable for statistical analysis.
- It can be used in conjunction with
mean()to comprehensively describe the distribution characteristics of the data. - In finance, standard deviation is often used to measure risk (volatility).
- The square of the standard deviation is the variance (
var())。
In actual data analysis, standard deviation is often used together with indicators such as mean and median to form a complete understanding of data distribution.
Other Extensions
Pandas Common Functions