Pandas Series.corr() Function

Pandas 常用函数Commonly Used Pandas Functions


Series.corr()is a function in Pandas used to calculate the correlation coefficient between two Series. The correlation coefficient measures the strength of the linear relationship between two variables, ranging from -1 to 1.

A correlation coefficient close to 1 indicates a positive correlation, close to -1 indicates a negative correlation, and close to 0 indicates almost no correlation. It is widely used in scenarios such as data analysis, feature selection, and regression analysis.


Basic Syntax and Parameters

corr()is a member function of the Series object, and requires another Series to be passed as a parameter.

Syntax Format

Series.corr(other, method='pearson', min_periods=None, **kwargs)

Parameter Description

Parameter Type Description Default Value
other Series The other Series used to calculate the correlation coefficient. Required
method str Method for calculating the correlation coefficient. Optional values: 'pearson' (Pearson), 'spearman' (Spearman), 'kendall' (Kendall). 'pearson'
min_periods int The minimum number of valid observations required for the calculation. If the number of valid values is less than this, NaN is returned. None

Return Value

  • Return Type:float
  • Description: Returns the correlation coefficient between the two Series, ranging from -1 to 1.

Examples

Let us, through a series of examples from simple to complex, thoroughly masterSeries.corr()the usage.

Example 1: Basic Usage - Pearson Correlation Coefficient

The Pearson correlation coefficient is the most commonly used correlation coefficient, measuring the linear relationship between two variables.

Example

import pandas as pd

# Create two Series: study hours and exam scores
study_hours = pd.Series([1, 2, 3, 4, 5, 6, 7, 8])
exam_scores = pd.Series([55, 60, 65, 70, 75, 80, 85, 90])

print("Study time (hours):")
print(study_hours)
print()

print("Exam score (points):")
print(exam_scores)
print()

# Calculate the Pearson correlation coefficient
correlation = study_hours.corr(exam_scores)

print(f"Pearson correlation coefficient: {correlation:.4f}")
print()
print("Analysis: The correlation coefficient is close to 1, indicating a strong positive correlation between study time and exam scores.")

Output:

学习时间(小时):
0    1
1    2
2    3
3    4
4    5
5    6
6    7
7    8
dtype: int64

考试成绩(分):
0    55
1    60
2    65
3    70
4    75
5    80
6    85
7    90
dtype: int64

皮尔逊相关系数:1.0000

Code explanation:

  • The correlation coefficient is 1.0, indicating a perfect positive correlation.
  • This means that for every additional hour of study time, the exam score increases by 5 points.

Example 2: Negative Correlation and No Correlation Examples

Shows different types of correlation.

Example

import pandas as pd

# Positively correlated data: age and income
ages = pd.Series([25, 30, 35, 40, 45, 50])
income = pd.Series([4000, 5000, 6000, 7000, 8000, 9000])

print("Positive correlation example - Age vs Income:")
print(f"Correlation coefficient: {ages.corr(income):.4f}")
print()

# Negatively correlated data: gaming time and academic performance
game_hours = pd.Series([1, 2, 3, 4, 5, 6])
grades = pd.Series([95, 90, 82, 75, 65, 55])

print("Negative correlation example - Gaming time vs Academic performance:")
print(f"Correlation coefficient: {game_hours.corr(grades):.4f}")
print()

# Uncorrelated data: shoe size and IQ
shoe_sizes = pd.Series([36, 37, 38, 39, 40, 41])
iq_scores = pd.Series([100, 105, 95, 110, 98, 102])

print("No correlation example - Shoe size vs IQ:")
print(f"Correlation coefficient: {shoe_sizes.corr(iq_scores):.4f}")
print()

print("Interpretation of correlation coefficient:")
print(" |r| > 0.7: Strong correlation")
print(" 0.4 < |r| <= 0.7: Moderate correlation")
print(" 0.2 < |r| <= 0.4: Weak correlation")
print(" |r| <= 0.2: Almost no correlation")

Output:

正相关示例 - 年龄与收入:
相关系数:1.0000
负相关示例 - 游戏时间与学习成绩:
相关系数:-0.9856
无相关示例 - 鞋码与智商:
相关系数:0.0857

Example 3: Using Different Correlation Coefficient Methods

Demonstrates the usage of Spearman and Kendall correlation coefficients.

Example

import pandas as pd

# Create data with a non-linear relationship
x = pd.Series([1, 2, 3, 4, 5, 6, 7, 8, 9, 10])
# Use an exponential relationship, not a purely linear one
y = pd.Series([1, 4, 9, 16, 25, 36, 49, 64, 81, 100])

print("Data X:")
print(x)
print()

print("Data Y (X squared):")
print(y)
print()

# Pearson correlation coefficient (measures linear relationship)
pearson_corr = x.corr(y, method='pearson')
print(f"Pearson correlation coefficient: {pearson_corr:.4f}")

# Spearman correlation coefficient (measures monotonic relationship)
spearman_corr = x.corr(y, method='spearman')
print(f"Spearman correlation coefficient: {spearman_corr:.4f}")

# Kendall correlation coefficient
kendall_corr = x.corr(y, method='kendall')
print(f"Kendall correlation coefficient: {kendall_corr:.4f}")
print()

print("Analysis:")
print("Although X and Y have a perfect monotonic relationship (y = x^2),")
print("But the Pearson correlation coefficient is not 1, because they are not linearly related.")
print("The Spearman and Kendall correlation coefficients are 1, because they measure monotonic relationships.")

Output:

数据 X:
0    1
1    2
2    3
3    4
4    5
5    6
6    7
7    8
8    9
9   10
dtype: int64

数据 Y(X 的平方):
0      1
1      4
2      9
3     16
4     25
5     36
6     49
7     64
8     81
9    100
dtype: int64

皮尔逊相关系数:0.9746
斯皮尔曼相关系数:1.0000
肯德尔相关系数:1.0000

Example 4: Using the min_periods Parameter

Controls the minimum number of valid observations required for the calculation.

Example

import pandas as pd
import numpy as np

# Create two Series containing missing values
s1 = pd.Series([1, 2, np.nan, 4, 5, np.nan, 7, 8])
s2 = pd.Series([2, 4, 6, 8, 10, 12, 14, 16])

print("Series 1:")
print(s1)
print()

print("Series 2:")
print(s2)
print()

# By default, at least 2 valid values are required
corr_default = s1.corr(s2)
print(f"Default (min_periods=None) correlation coefficient: {corr_default:.4f}")

# Set a minimum of 6 valid values
corr_6 = s1.corr(s2, min_periods=6)
print(f"min_periods=6 correlation coefficient: {corr_6:.4f}")

# Set a minimum of 7 valid values (only 6 valid values, returns NaN)
corr_7 = s1.corr(s2, min_periods=7)
print(f"min_periods=7 correlation coefficient: {corr_7}")

Output:

Series 1:
0    1.0
1    2.0
2       NaN
3    4.0
4    5.0
5       NaN
6    7.0
7    8.0
dtype: float64

Series 2:
0     2
1     4
2     6
3     8
4    10
5    12
6    14
7    16
dtype: int64

默认(min_periods=None)相关系数:1.0000
min_periods=6 相关系数:1.0000
min_periods=7 相关系数:nan

Example 5: Practical Application - Analyzing Economic Indicator Correlations

Demonstrates the application of correlation coefficients in real-world data analysis.

Example

import pandas as pd

# Create simulated macroeconomic data
gdp_growth = pd.Series([6.5, 7.2, 8.1, 7.8, 6.9, 5.5, 4.8])
unemployment = pd.Series([4.2, 3.8, 3.5, 3.7, 4.0, 4.8, 5.2])
inflation = pd.Series([2.1, 2.5, 3.2, 3.0, 2.4, 1.8, 1.5])
stock_index = pd.Series([3000, 3200, 3500, 3400, 3100, 2800, 2600])

print("Macroeconomic data:")
print(f"GDP growth rate: {gdp_growth.values}")
print(f"Unemployment rate: {unemployment.values}")
print(f"Inflation rate: {inflation.values}")
print(f"Stock index: {stock_index.values}")
print()

# Analyze the correlation between indicators
print("Correlation analysis:")
print(f"GDP growth rate vs unemployment rate: {gdp_growth.corr(unemployment):.4f}")
print(f"GDP growth rate vs stock index: {gdp_growth.corr(stock_index):.4f}")
print(f"Unemployment rate vs stock index: {unemployment.corr(stock_index):.4f}")
print(f"Inflation rate vs stock index: {inflation.corr(stock_index):.4f}")
print()

print("Conclusion:")
print("- GDP growth rate is negatively correlated with unemployment rate (when the economy is good, unemployment is low)")
print("- GDP growth rate is strongly positively correlated with stock index (when the economy is good, the stock market rises)")
print("- Unemployment rate is negatively correlated with stock index")

Output:

宏观经济数据:
GDP 增长率:[6.5 7.2 8.1 7.8 6.9 5.5 4.8]
失业率:[4.2 3.8 3.5 3.7 4.  4.  5.2]
通胀率:[2.1 2.5 3.2 3.  2.4 1.8 1.5]
股指:[3000 3200 3500 3400 3100 2800 2600]

相关性分析:
GDP 增长率 vs 失业率:-0.9036
GDP 增长率 vs 股指:0.9434
失业率 vs 股指:-0.9712
通胀率 vs 股指:-0.9712

Code explanation:

  • The correlation coefficient between GDP growth rate and unemployment rate is -0.9036, indicating a strong negative correlation.
  • The correlation coefficient between GDP growth rate and stock index is 0.9434, indicating a strong positive correlation.
  • These correlations are consistent with economic intuition.

Notes

  • The correlation coefficient can only measure linear relationships and may be inaccurate for non-linear relationships.
  • A correlation coefficient close to 0 does not mean there is no relationship; there may be a non-linear relationship.
  • Correlation does not equal causation; it needs to be judged in combination with business logic.
  • The Spearman correlation coefficient is suitable for ranked data or non-linear monotonic relationships.
  • The Kendall correlation coefficient is more robust to noise in data, but is slower to compute.

Summary

Series.corr()It is a powerful tool for analyzing the relationship between two variables. Its main features include:

  • Supports three correlation coefficient calculation methods: Pearson, Spearman, and Kendall.
  • The Pearson correlation coefficient is suitable for measuring linear relationships.
  • The Spearman and Kendall correlation coefficients are suitable for measuring monotonic relationships.
  • You can control the minimum number of valid observations.

In practical applications, correlation coefficients are often used for feature selection, data exploration, and hypothesis testing. However, note that correlation analysis can only reveal statistical associations and cannot directly infer causal relationships.

Pandas 常用函数Pandas Common Functions

Other Extensions