Regression Model Evaluation

In the world of machine learning, building a model is only the first step. Just as we bake a cake, we must taste it to know whether it is successful. For regression models (a type of model used to predict continuous numerical values, such as house prices, temperature, or sales),model evaluationis ourtastingstep. It tells us how accurate the model's predictions are, where it does well, and where there is room for improvement.

This article will guide you through a systematic study of regression model evaluation methods, from core concepts to specific metrics, and then to code practice, enabling you not only to understand evaluation reports but also to "score" your model yourself.


Core Concepts of Regression Models and Evaluation

Before diving into metrics, we need to understand a few fundamental concepts; they are the cornerstone of all evaluation methods.

What is a regression problem?

Regression is a type of supervised learning whose goal is to predict acontinuous numerical value。

  • Example: Predict the selling price (a continuous numerical value) based on the house's area, age, and location (features).

Why do we need evaluation?

  1. Model selection: Compare which of different algorithms (e.g., linear regression, decision tree regression) performs better on your data.
  2. Parameter tuning: When adjusting model parameters (e.g., regularization strength), we need an objective criterion to judge whether the adjustment is effective.
  3. Performance confirmation: Ensure the model has reliable predictive ability on unseen data, avoiding "armchair strategy".

Key Term: Error

All evaluation metrics revolve around "error". Error is the difference between the model's predicted value (ŷ) and the true value (y).

  • Error at a single point:误差 = y - ŷ
  • The essence of evaluation issummarizing and analyzing these errors in different ways.。

Detailed Explanation of Core Evaluation Metrics

We will introduce the most commonly used and most important evaluation metrics, using a simple example throughout.

Suppose we predicted the prices of 5 houses:

True price (y) Predicted price (ŷ) Error (y - ŷ)
200 210 -10
150 145 5
300 310 -10
400 380 20
250 255 -5

1. Mean Absolute Error (MAE)

Intuitive understanding: Add up the absolute values of all prediction "errors" and then compute the average. It directly reflects "on average, by how many units the predicted value deviates from the true value".

Calculation formula: MAE = (1/n) * Σ|y_i - ŷ_i|wherenis the number of samples,Σis the summation symbol.

Calculation example: MAE = (|-10| + |5| + |-10| + |20| + |-5|) / 5 = (10+5+10+20+5)/5 = 50/5 = 10

Interpretation: On average, our house price prediction deviates from the true price by100,000 yuan。

Characteristics:

  • Advantages: Intuitive and easy to understand; not excessively affected by extreme error values (outliers).
  • Disadvantages: The absolute value function is not differentiable everywhere in mathematics, which is inconvenient in some optimization scenarios.

2. Mean Squared Error (MSE)

Intuitive understanding: First "square" each error (to eliminate the negative sign and amplify errors), then take the average. It isvery sensitive to large errors.。

Calculation formula: MSE = (1/n) * Σ(y_i - ŷ_i)^2

Calculation example: MSE = [(-10)^2 + (5)^2 + (-10)^2 + (20)^2 + (-5)^2] / 5 = (100+25+100+400+25)/5 = 650/5 = 130

Interpretation: The average of the squared errors is 130. This value itself has no direct unit meaning (it is 10,000 yuan squared), but it is very effective for comparing models.

Characteristics:

  • Advantages: Excellent mathematical properties (differentiable everywhere); it is the objective function minimized during training of many models (e.g., linear regression).
  • Disadvantages: Its dimension differs from the original data, making the magnitude difficult to interpret; sensitive to outliers.

3. Root Mean Squared Error (RMSE)

Intuitive understanding: It is the square root of MSE. It is equivalent to "restoring" MSE to its original form, bringing it back to the same dimension as the true values and predicted values.

Calculation formula: RMSE = sqrt(MSE)

Calculation example: RMSE = sqrt(130) ≈ 11.4

Interpretation: On average, our predictions deviate from the true values by approximately114,000 yuan. This interpretation is similar to MAE (100,000 yuan), but because RMSE squares first and then takes the square root, it gives larger errors higher weight, so it is usually larger than MAE.

Characteristics:

  • Advantages: It has the same dimension as the original data, with stronger interpretability than MSE; it punishes large errors more heavily and is commonly used in scenarios where large errors matter (e.g., financial risk prediction).
  • Disadvantages: It is also sensitive to outliers.

4. R² Score (Coefficient of Determination)

Intuitive understanding: How much better is my model than "guessing blindly" (guessing with the mean)? It is aratio value, measuring the model's ability to explain the variation in the data.

Calculation formula: R² = 1 - (Σ(y_i - ŷ_i)^2 / Σ(y_i - y_mean)^2)wherey_meanis the mean of the true values.

"Blind-guessing" model: For any house, I predict the mean house price (y_mean)。

  • calculated asy_mean = (200+150+300+400+250)/5 = 260
  • The sum of squared errors of the "blind-guessing" model:Σ(y_i - 260)^2 = 1600+12100+1600+19600+100 = 35000

Calculation example: R² = 1 - (650 / 35000) = 1 - 0.01857 ≈ 0.9814

Interpretation: Our model can explain the target variable (house price)98.14%fluctuation. This indicates that the model fits very well.

Characteristics:

  • Range: Theoretically, the range of R² is(-∞, 1]。
    • R² = 1: Perfect prediction.
    • R² = 0: The model's effect is the same as directly predicting with the mean.
    • R² < 0: The model is even worse than directly predicting with the mean (indicating the model is completely unsuitable for the data).
  • Advantages: Dimensionless, easy to compare model performance across different datasets.
  • Disadvantages: As model features increase, R² naturally increases even if the added features are useless, which may lead to overfitting.

Metric Comparison and How to Choose

For a clearer understanding, let's summarize with a table:

Metric Calculation formula (simplified) Dimension Outlier sensitivity Characteristics and applicable scenarios
MAE Mean Absolute Error Same as y Insensitive Most intuitive interpretation. Focuses on "how large the average deviation is"; suitable for all regression scenarios, especially data with many outliers.
MSE Mean Squared Error Square of y Very sensitive Good mathematical properties. It is a common loss function for model training. The value itself is not easy to interpret.
RMSE Square root of MSE Same as y Very sensitive Most commonly used. It combines the mathematical advantages of MSE with interpretability. It heavily penalizes large errors and is popular in finance, forecasting, and other fields.
R² 1 - (model error / baseline error) Dimensionless Sensitive Relative performance metric. Used to compare the improvement of a model relative to a simple baseline. Easy to compare across datasets.

Selection recommendations:

  1. Prefer reporting RMSE and R²: RMSE gives the actual magnitude of the error, while R² gives the relative performance of the model. This is the most versatile combination.
  2. When there are many outliers in the data, pay attention to MAE。
  3. Use MSE as the loss function during model training and optimization.。
  4. Never look at only one metric! Combine multiple metrics to comprehensively evaluate the model.

Code Practice: Evaluation with Scikit-learn

Now, let's put it into practice with Python and the machine learning library Scikit-learn.

Environment Setup

Make sure the necessary libraries are installed:

Example

pip install numpy scikit-learn matplotlib

Complete Example Code

Example

import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score

# 1. Create simulated data
np.random.seed(42) # Ensure consistent results each run
X = 2 * np.random.rand(100, 1) # 100 samples, 1 feature, range [0,2)
y = 4 + 3 * X + np.random.randn(100, 1) # True relationship: y = 4 + 3x + noise

# 2. Split into training and test sets (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# 3. Train a simple linear regression model
model = LinearRegression()
model.fit(X_train, y_train)

# 4. Make predictions on the test set
y_pred = model.predict(X_test)

# 5. Calculate all evaluation metrics
mae = mean_absolute_error(y_test, y_pred)
mse = mean_squared_error(y_test, y_pred)
rmse = np.sqrt(mse) # Or use mean_squared_error(y_test, y_pred, squared=False)
r2 = r2_score(y_test, y_pred)

# 6. Print evaluation results
print("=== Regression Model Evaluation Report ===")
print(f"Mean Absolute Error (MAE): {mae:.4f}")
print(f"Mean Squared Error (MSE): {mse:.4f}")
print(f"Root Mean Squared Error (RMSE): {rmse:.4f}")
print(f"Coefficient of Determination (R² Score): {r2:.4f}")
print("\nModel coefficients:)
print(f" Intercept: {model.intercept_)
print(f" Slope (Coefficient for X): {model.coef_)

# 7. Visualize the results
plt.figure(figsize=(10, 5))

# Subplot 1: Scatter plot of true values vs predicted values
plt.subplot(1, 2, 1)
plt.scatter(y_test, y_pred, alpha=0.7, edgecolors='k')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--', lw=2, label='Perfect prediction line')
plt.xlabel('True values (y_test)')
plt.ylabel('Predicted values (y_pred)')
plt.title('True values vs Predicted values')
plt.legend()
plt.grid(True, linestyle='--', alpha=0.7)

# Subplot 2: Histogram of error distribution
plt.subplot(1, 2, 2)
errors = y_test.flatten() - y_pred.flatten()
plt.hist(errors, bins=15, edgecolor='black', alpha=0.7)
plt.axvline(x=0, color='r', linestyle='--', label='Zero error line')
plt.xlabel('Prediction error')
plt.ylabel('Frequency')
plt.title('Prediction error distribution')
plt.legend()
plt.grid(True, linestyle='--', alpha=0.7, axis='y')

plt.tight_layout()
plt.show()

Line-by-Line Code Explanation

  1. Create data: we simulatedy = 4 + 3x + 噪声data, which is a typical linear relationship plus random noise.
  2. Split the dataset:train_test_splitRandomly split the data into two parts. Use the training set (X_train, y_train) to learn the model, and use the unseen test set (X_test, y_test) to evaluate its generalization ability. This is a crucial step in evaluation!
  3. Train the model: CreateLinearRegressionobject and call.fit()method to train.
  4. Make predictions: Call.predict()method on the test set featuresX_testto generate predicted valuesy_pred。
  5. Calculate metrics: Use Scikit-learn'smetricsfunctions in the module to easily compute MAE, MSE, R². RMSE is obtained by taking the square root of MSE.
  6. Print results: Format and output all metrics and the parameters learned by the model.
  7. Visualization:
    • Left plot: Ideally, the scatter points should closely align around the red dashed line (perfect prediction line).
    • Right plot: The error distribution histogram should be roughly centered at 0 and approximately normally distributed, which is usually a good sign.

Hands-on Exercises and Summary

Hands-on Exercise

  1. Run the code: Copy the code above into your Python environment and run it. Observe the output results and plots.
  2. Modify the data: Try increasingnp.random.randn()noise scale in it, then rerun. Observe how all evaluation metrics change? How do the plots differ?
  3. Change the model: Try usingfrom sklearn.tree import DecisionTreeRegressorto replace the linear regression model. Compare the evaluation metrics of the two models.
  4. Calculation exercise: Manually calculate the MAE, MSE, RMSE, and R² for the following three samples.
    • True values y: [10, 20, 30]
    • Predicted values ŷ: [12, 18, 28]
Other extensions