Pandas Series.corr() Function
Commonly Used Pandas Functions
Series.corr()is a function in Pandas used to calculate the correlation coefficient between two Series. The correlation coefficient measures the strength of the linear relationship between two variables, ranging from -1 to 1.
A correlation coefficient close to 1 indicates a positive correlation, close to -1 indicates a negative correlation, and close to 0 indicates almost no correlation. It is widely used in scenarios such as data analysis, feature selection, and regression analysis.
Basic Syntax and Parameters
corr()is a member function of the Series object, and requires another Series to be passed as a parameter.
Syntax Format
Series.corr(other, method='pearson', min_periods=None, **kwargs)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| other | Series | The other Series used to calculate the correlation coefficient. | Required |
| method | str | Method for calculating the correlation coefficient. Optional values: 'pearson' (Pearson), 'spearman' (Spearman), 'kendall' (Kendall). | 'pearson' |
| min_periods | int | The minimum number of valid observations required for the calculation. If the number of valid values is less than this, NaN is returned. | None |
Return Value
- Return Type:
float - Description: Returns the correlation coefficient between the two Series, ranging from -1 to 1.
Examples
Let us, through a series of examples from simple to complex, thoroughly masterSeries.corr()the usage.
Example 1: Basic Usage - Pearson Correlation Coefficient
The Pearson correlation coefficient is the most commonly used correlation coefficient, measuring the linear relationship between two variables.
Example
# Create two Series: study hours and exam scores
study_hours = pd.Series([1, 2, 3, 4, 5, 6, 7, 8])
exam_scores = pd.Series([55, 60, 65, 70, 75, 80, 85, 90])
print("Study time (hours):")
print(study_hours)
print()
print("Exam score (points):")
print(exam_scores)
print()
# Calculate the Pearson correlation coefficient
correlation = study_hours.corr(exam_scores)
print(f"Pearson correlation coefficient: {correlation:.4f}")
print()
print("Analysis: The correlation coefficient is close to 1, indicating a strong positive correlation between study time and exam scores.")
Output:
学习时间(小时): 0 1 1 2 2 3 3 4 4 5 5 6 6 7 7 8 dtype: int64 考试成绩(分): 0 55 1 60 2 65 3 70 4 75 5 80 6 85 7 90 dtype: int64 皮尔逊相关系数:1.0000
Code explanation:
- The correlation coefficient is 1.0, indicating a perfect positive correlation.
- This means that for every additional hour of study time, the exam score increases by 5 points.
Example 2: Negative Correlation and No Correlation Examples
Shows different types of correlation.
Example
# Positively correlated data: age and income
ages = pd.Series([25, 30, 35, 40, 45, 50])
income = pd.Series([4000, 5000, 6000, 7000, 8000, 9000])
print("Positive correlation example - Age vs Income:")
print(f"Correlation coefficient: {ages.corr(income):.4f}")
print()
# Negatively correlated data: gaming time and academic performance
game_hours = pd.Series([1, 2, 3, 4, 5, 6])
grades = pd.Series([95, 90, 82, 75, 65, 55])
print("Negative correlation example - Gaming time vs Academic performance:")
print(f"Correlation coefficient: {game_hours.corr(grades):.4f}")
print()
# Uncorrelated data: shoe size and IQ
shoe_sizes = pd.Series([36, 37, 38, 39, 40, 41])
iq_scores = pd.Series([100, 105, 95, 110, 98, 102])
print("No correlation example - Shoe size vs IQ:")
print(f"Correlation coefficient: {shoe_sizes.corr(iq_scores):.4f}")
print()
print("Interpretation of correlation coefficient:")
print(" |r| > 0.7: Strong correlation")
print(" 0.4 < |r| <= 0.7: Moderate correlation")
print(" 0.2 < |r| <= 0.4: Weak correlation")
print(" |r| <= 0.2: Almost no correlation")
Output:
正相关示例 - 年龄与收入: 相关系数:1.0000 负相关示例 - 游戏时间与学习成绩: 相关系数:-0.9856 无相关示例 - 鞋码与智商: 相关系数:0.0857
Example 3: Using Different Correlation Coefficient Methods
Demonstrates the usage of Spearman and Kendall correlation coefficients.
Example
# Create data with a non-linear relationship
x = pd.Series([1, 2, 3, 4, 5, 6, 7, 8, 9, 10])
# Use an exponential relationship, not a purely linear one
y = pd.Series([1, 4, 9, 16, 25, 36, 49, 64, 81, 100])
print("Data X:")
print(x)
print()
print("Data Y (X squared):")
print(y)
print()
# Pearson correlation coefficient (measures linear relationship)
pearson_corr = x.corr(y, method='pearson')
print(f"Pearson correlation coefficient: {pearson_corr:.4f}")
# Spearman correlation coefficient (measures monotonic relationship)
spearman_corr = x.corr(y, method='spearman')
print(f"Spearman correlation coefficient: {spearman_corr:.4f}")
# Kendall correlation coefficient
kendall_corr = x.corr(y, method='kendall')
print(f"Kendall correlation coefficient: {kendall_corr:.4f}")
print()
print("Analysis:")
print("Although X and Y have a perfect monotonic relationship (y = x^2),")
print("But the Pearson correlation coefficient is not 1, because they are not linearly related.")
print("The Spearman and Kendall correlation coefficients are 1, because they measure monotonic relationships.")
Output:
数据 X: 0 1 1 2 2 3 3 4 4 5 5 6 6 7 7 8 8 9 9 10 dtype: int64 数据 Y(X 的平方): 0 1 1 4 2 9 3 16 4 25 5 36 6 49 7 64 8 81 9 100 dtype: int64 皮尔逊相关系数:0.9746 斯皮尔曼相关系数:1.0000 肯德尔相关系数:1.0000
Example 4: Using the min_periods Parameter
Controls the minimum number of valid observations required for the calculation.
Example
import numpy as np
# Create two Series containing missing values
s1 = pd.Series([1, 2, np.nan, 4, 5, np.nan, 7, 8])
s2 = pd.Series([2, 4, 6, 8, 10, 12, 14, 16])
print("Series 1:")
print(s1)
print()
print("Series 2:")
print(s2)
print()
# By default, at least 2 valid values are required
corr_default = s1.corr(s2)
print(f"Default (min_periods=None) correlation coefficient: {corr_default:.4f}")
# Set a minimum of 6 valid values
corr_6 = s1.corr(s2, min_periods=6)
print(f"min_periods=6 correlation coefficient: {corr_6:.4f}")
# Set a minimum of 7 valid values (only 6 valid values, returns NaN)
corr_7 = s1.corr(s2, min_periods=7)
print(f"min_periods=7 correlation coefficient: {corr_7}")
Output:
Series 1: 0 1.0 1 2.0 2 NaN 3 4.0 4 5.0 5 NaN 6 7.0 7 8.0 dtype: float64 Series 2: 0 2 1 4 2 6 3 8 4 10 5 12 6 14 7 16 dtype: int64 默认(min_periods=None)相关系数:1.0000 min_periods=6 相关系数:1.0000 min_periods=7 相关系数:nan
Example 5: Practical Application - Analyzing Economic Indicator Correlations
Demonstrates the application of correlation coefficients in real-world data analysis.
Example
# Create simulated macroeconomic data
gdp_growth = pd.Series([6.5, 7.2, 8.1, 7.8, 6.9, 5.5, 4.8])
unemployment = pd.Series([4.2, 3.8, 3.5, 3.7, 4.0, 4.8, 5.2])
inflation = pd.Series([2.1, 2.5, 3.2, 3.0, 2.4, 1.8, 1.5])
stock_index = pd.Series([3000, 3200, 3500, 3400, 3100, 2800, 2600])
print("Macroeconomic data:")
print(f"GDP growth rate: {gdp_growth.values}")
print(f"Unemployment rate: {unemployment.values}")
print(f"Inflation rate: {inflation.values}")
print(f"Stock index: {stock_index.values}")
print()
# Analyze the correlation between indicators
print("Correlation analysis:")
print(f"GDP growth rate vs unemployment rate: {gdp_growth.corr(unemployment):.4f}")
print(f"GDP growth rate vs stock index: {gdp_growth.corr(stock_index):.4f}")
print(f"Unemployment rate vs stock index: {unemployment.corr(stock_index):.4f}")
print(f"Inflation rate vs stock index: {inflation.corr(stock_index):.4f}")
print()
print("Conclusion:")
print("- GDP growth rate is negatively correlated with unemployment rate (when the economy is good, unemployment is low)")
print("- GDP growth rate is strongly positively correlated with stock index (when the economy is good, the stock market rises)")
print("- Unemployment rate is negatively correlated with stock index")
Output:
宏观经济数据: GDP 增长率:[6.5 7.2 8.1 7.8 6.9 5.5 4.8] 失业率:[4.2 3.8 3.5 3.7 4. 4. 5.2] 通胀率:[2.1 2.5 3.2 3. 2.4 1.8 1.5] 股指:[3000 3200 3500 3400 3100 2800 2600] 相关性分析: GDP 增长率 vs 失业率:-0.9036 GDP 增长率 vs 股指:0.9434 失业率 vs 股指:-0.9712 通胀率 vs 股指:-0.9712
Code explanation:
- The correlation coefficient between GDP growth rate and unemployment rate is -0.9036, indicating a strong negative correlation.
- The correlation coefficient between GDP growth rate and stock index is 0.9434, indicating a strong positive correlation.
- These correlations are consistent with economic intuition.
Notes
- The correlation coefficient can only measure linear relationships and may be inaccurate for non-linear relationships.
- A correlation coefficient close to 0 does not mean there is no relationship; there may be a non-linear relationship.
- Correlation does not equal causation; it needs to be judged in combination with business logic.
- The Spearman correlation coefficient is suitable for ranked data or non-linear monotonic relationships.
- The Kendall correlation coefficient is more robust to noise in data, but is slower to compute.
Summary
Series.corr()It is a powerful tool for analyzing the relationship between two variables. Its main features include:
- Supports three correlation coefficient calculation methods: Pearson, Spearman, and Kendall.
- The Pearson correlation coefficient is suitable for measuring linear relationships.
- The Spearman and Kendall correlation coefficients are suitable for measuring monotonic relationships.
- You can control the minimum number of valid observations.
In practical applications, correlation coefficients are often used for feature selection, data exploration, and hypothesis testing. However, note that correlation analysis can only reveal statistical associations and cannot directly infer causal relationships.
Other Extensions