Multiple Linear Regression

In the previous article, we explored simple linear regression, which helped us understand how one feature (independent variable) affects a target (dependent variable). However, problems in the real world are often more complex. For example, when predicting house prices, we cannot only look at the house area; we also need to consider multiple factors such as the number of bedrooms, geographic location, and house age. At this point, we need:Multiple Linear Regression。

In simple terms, multiple linear regression is a natural extension of simple linear regression. It allows us to simultaneously analyze the influence ofmultiple independent variableson a dependent variable. This article will take you from zero to a comprehensive understanding of the core concepts, mathematical principles, implementation methods, and practical applications of multiple linear regression.


1. What is Multiple Linear Regression?

1.1 Core Concepts

Multiple linear regression is a statistical method used to establish a linear relationship betweenmultiple independent variables(also called features, explanatory variables) andone continuous dependent variable(also called target, response variable).

A vivid metaphor:Imagine you are a chef, adjusting the taste of a soup (the target). Simple linear regression is like adjusting the saltiness using only the amount of salt. Multiple linear regression is like controlling the amounts of salt, sugar, chili, soy sauce, and other seasonings (features) at the same time to comprehensively determine the final taste of the soup. Multiple regression allows you to analyze the specific contribution of each seasoning to the taste.

1.2 Model Formula

The mathematical expression of the multiple linear regression model is as follows:

y = β₀ + β₁x₁ + β₂x₂ + ... + βₙxₙ + ε

Let's break down each part of this formula:

Symbol Name Meaning
y Dependent variable The target value we want to predict (e.g., house price).
x₁, x₂, ..., xₙ Independent variable Features used to predicty(e.g., area, number of bedrooms, house age).
β₀ Intercept term When all independent variables are 0,ythe baseline predicted value.
β₁, β₂, ..., βₙ Regression coefficient The weight of each independent variablexᵢ. It indicates:When other features remain unchanged,xᵢfor each additional 1 unit of that variable,ythe average change in the targetβᵢis that many units.This is the core of multiple regression analysis.
ε Error term Random fluctuation that the model cannot explain (e.g., measurement error, unknown factors).

Example of interpreting the formula:Suppose we predict house price (y), the model is:房价 = 50000 + 3000 * 面积 + 10000 * 卧室数 - 2000 * 房龄

  • β₀ = 50000β₀: Theoretically, the baseline price of a "house" with area 0, bedrooms 0, and house age 0.
  • β₁ = 3000β₁: With the number of bedrooms and house age unchanged, for every additional 1 square meter of area, the house price increases by an average of 3000 yuan.
  • β₂ = 10000β₂: With area and house age unchanged, one more bedroom increases the house price by an average of 10000 yuan.
  • β₃ = -2000β₃: With area and number of bedrooms unchanged, for each additional year of house age, the house price decreases by an average of 2000 yuan.

2. How to "train" a multiple linear regression model?

"Training" a model essentially means finding a set of optimal regression coefficients based on our existing data,(β₀, β₁, ..., βₙ)so that the model's predicted values are as close as possible to the true values. This process is usually completed throughthe least squares method.

2.1 Objective: Minimize the loss function

We use theresidual sum of squares (RSS)as the loss function to measure the model's prediction error.RSS = Σ(yᵢ - ŷᵢ)²where,yᵢis the i-thitrue value,ŷᵢis the model's prediction for the i-thisample.

The training objective is to find a set of coefficients that minimizes the value of RSS.

2.2 Solution process (matrix form)

When the number of features is large, using matrix operations can represent and solve the model more efficiently. The model can be written as:Y = Xβ + εwhere:

  • Yy is a column vector containing all target values.
  • XX is the design matrix. The first column is usually all 1s (corresponding to the intercept termβ₀), and each subsequent column corresponds to a feature.
  • ββ is a column vector containing all regression coefficients.
  • εε is the error term vector.

Through least squares derivation, the optimal solution (closed-form solution) for the coefficientsβcan be obtained as:β = (XᵀX)⁻¹XᵀYThis formula is theoretically perfect, but in actual programming, we usually use numerical optimization libraries (such asscikit-learn) to compute efficiently and stably. It automatically handles complex operations such as matrix inversion.


3. Implementing Multiple Linear Regression in Python

Let's use a popular machine learning library through a complete examplescikit-learnto build a multiple linear regression model.

3.1 Environment Setup and Data Loading

First, make sure the necessary libraries are installed, and load a sample dataset.

Example

# Import necessary libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score
from sklearn.datasets import fetch_california_housing # A classic multivariate dataset

# Load the California housing dataset
california = fetch_california_housing()
df = pd.DataFrame(california.data, columns=california.feature_names)
df['MedHouseVal'] = california.target # Add target column: median house value

print("Dataset shape:", df.shape)
print("\n"First 5 rows of data:")
print(df.head())
print("\n"Feature description:")
print(california.DESCR[:500]) # Print part of the description

3.2 Data Exploration and Preprocessing

Before modeling, it is crucial to understand the basic situation of the data.

Example

# 1. View basic information about the data
print(df.info())
print("\n"Basic statistical description:")
print(df.describe())

# 2. Split into features (X) and target (y)
X = df.drop('MedHouseVal', axis=1) # Feature matrix: contains all columns except house price
y = df['MedHouseVal'] # Target vector: house price

# 3. Split into training set and test set (70% training, 30% test)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print(f"\n"Training set samples: {X_train.shape)

3.3 Create, Train, and Evaluate the Model

Now, we create a linear regression model, "fit" it with the training data, and evaluate its performance on the test set.

Example

# 1. Create model instance
model = LinearRegression()

# 2. Train the model (fit the data)
model.fit(X_train, y_train)

# 3. Use the trained model to make predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)

# 4. Evaluate model performance
# Mean Squared Error (MSE) - lower is better
train_mse = mean_squared_error(y_train, y_train_pred)
test_mse = mean_squared_error(y_test, y_test_pred)

# Coefficient of determination (R²) - closer to 1 is better, indicating stronger explanatory power
train_r2 = r2_score(y_train, y_train_pred)
test_r2 = r2_score(y_test, y_test_pred)

print("=== Model Performance Evaluation ===")
print(f"Training set MSE: {train_mse:.4f}, R²: {train_r2:.4f}")
print(f"Test set MSE: {test_mse:.4f}, R²: {test_r2:.4f}")

# 5. View the learned model parameters
print("\n"=== Model Parameters ===")
print(f"Intercept (β₀): {model.intercept_:.4f}")
print("Regression coefficients (β₁, β₂, ...):")
for feature, coef in zip(X.columns, model.coef_):
    print(f"  {feature}: {coef:.4f}")

3.4 Interpreting Results and Visualization

Understand the meaning of coefficients and evaluation metrics.

Example

# Visualization: true values vs predicted values (test set)
plt.figure(figsize=(8, 6))
plt.scatter(y_test, y_test_pred, alpha=0.5)
plt.plot([y.min(), y.max()], [y.min(), y.max()], 'r--', lw=2) # Plot the ideal diagonal line
plt.xlabel('True house price')
plt.ylabel('Predicted house price')
plt.title('Multiple Linear Regression: True Values vs Predicted Values (Test Set)')
plt.grid(True, linestyle='--', alpha=0.7)
plt.show()

# Visualization: feature importance (approximately represented by the absolute value of coefficients)
features = X.columns
coefs = model.coef_
plt.figure(figsize=(10, 6))
bars = plt.barh(features, np.abs(coefs)) # Use absolute values to compare the magnitude of influence
plt.xlabel('Absolute value of regression coefficient')
plt.title('Magnitude of feature influence on house price (based on absolute coefficients)')
# Add value labels to the bar chart
for bar, coef in zip(bars, coefs):
    width = bar.get_width()
    plt.text(width + 0.01, bar.get_y() + bar.get_height()/2, f'{coef:.4f}', va='center')
plt.grid(True, axis='x', linestyle='--', alpha=0.7)
plt.tight_layout()
plt.show()

4. Considerations and Challenges of Multiple Linear Regression

4.1 Key Assumptions

The validity of the linear regression model rests on several statistical assumptions:

  1. Linearity: There is a linear relationship between the independent and dependent variables.
  2. Independence: Observations are independent of each other.
  3. Homoscedasticity: The variance of the error term remains constant across all observations.
  4. Normality: The error term follows a normal distribution.
  5. No multicollinearity: There should be no high correlation among independent variables.

4.2 Common Challenges and Solutions

Challenge Description Possible Consequences Solutions
Multicollinearity Features are highly correlated with each other. Coefficient estimates are unstable, making it difficult to interpret the effect of individual features. 1. Use a correlation matrix to check and remove highly correlated features.
2. Use Principal Component Analysis (PCA) for dimensionality reduction.
3. Use regularization methods (such as ridge regression).
Overfitting The model is too complex, perfectly fits the noise in the training data, and performs poorly on the test set. The test set error is much larger than the training set error. 1. Collect more data.
2. Use fewer features (feature selection).
3. Use regularization.
Nonlinear relationship The data relationship is inherently nonlinear. The model has poor predictive ability and a low R² value. 1. Transform features (e.g., polynomial features, logarithmic transformation).
2. Use nonlinear models (such as decision trees, neural networks).

5. Hands-on Practice: Solidify Your Understanding

Now it's your turn to practice! Complete the following exercises in order to consolidate your understanding of multiple linear regression.

Exercise 1: Model Interpretation

Using the model coefficients output by the code above, answer the following questions:

  1. Which feature has the greatestpositive impacton California housing prices? (Based on coefficient values)
  2. Which feature has anegative impact on California housing prices??
  3. How do you interpretAveRoomsthe coefficient of (average number of rooms)? Please describe in one complete sentence.

Exercise 2: Diagnosing Multicollinearity

  1. Compute the features'Xcorrelation matrix (df.corr())。
  2. Find the feature pairs whose absolute correlation is greater than 0.7. They may have multicollinearity.
  3. (Optional) Try removing one of the highly correlated features from the model, retrain, and observe how the coefficients and the R² score change.

Exercise 3: Try a New Dataset

  1. fromsklearn.datasetsload another regression dataset, such asload_diabetes(the diabetes dataset).
  2. Repeat the modeling steps from this article: data exploration, splitting, training, evaluation, visualization.
  3. Compare the model performance on the two datasets, and think about possible reasons.
Other Extensions