Cross-Validation
In the practice of machine learning, we often face a core problem: how to evaluate the quality of a model? You might think of using part of the data to train the model, and then using another part of unseen data to test its performance. This idea is completely correct, but how can we evaluate the model more reliably and stably? This isCross-Validationthe core problem to be solved.
In simple terms, cross-validation is a statistical method that evaluates a model's generalization ability (i.e., its ability to handle new data) by repeatedly splitting the dataset. It is like arranging a mock exam for the model, using multiple different mock test papers (data subsets) to check its true level, avoiding misjudgment due to the randomness of a single exam.
This article will guide you to deeply understand the principles, common methods, and key role of cross-validation in model optimization and engineering.
Why Do We Need Cross-Validation?
Before diving into technical details, let us first understand its necessity through an analogy.
Imagine you are a student about to take an important math exam. There are two methods to evaluate your level:
- Method A (Simple Split): The teacher randomly picks 10 questions from the question bank for you to take a mock exam, and then uses this score to predict your final exam score.
- Method B (Cross-Validation): The teacher divides the question bank into 5 parts. The first time, use parts 2, 3, 4, 5 to train you, and use part 1 for testing; the second time, use parts 1, 3, 4, 5 for training, and use part 2 for testing... Repeat this 5 times. Finally, take the average of the 5 test scores to evaluate you.
Which method is more reliable? ObviouslyMethod B。
- Method AThe risk of Method A is: if the 10 questions picked happen to be the types you are good at, your mock exam score will be inflated, leading to over-optimism about your true level; conversely, if the questions picked are all your blind spots, the score will be too low, leading to over-pessimism. The evaluation results fluctuate greatly and are unstable.
- Method BThrough multiple different training/testing combinations, it lets you experience various types of questions in the question bank, and the average score obtained is more representative of your overall, stable level, and predicts the final exam result more accurately.
In machine learning:
- the question bankis ourentire dataset。
- the studentis the model we want to trainmachine learning model。
- the mock exam scoreis the model'sevaluation metric(such as accuracy, mean squared error, etc.).
- the final examis the model's futurereal, unseen dataperformance on.
The core goal of cross-validation is to provide, for the model's generalization ability, amore robust and more unbiased estimate, thereby helping us perform more reliable model selection, parameter tuning, and performance evaluation.
Common Methods of Cross-Validation
Cross-validation has multiple implementation methods, suitable for different data and scenarios. The most common ones are introduced below.
1. Hold-Out Validation
This is the simplest and most intuitive method.
Example
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
# Construct runnable data
# 200 samples, 4 features, binary classification
np.random.seed(42)
X = np.random.randn(200, 4)
y = (X[:, 0] + X[:, 1] * 0.7 - X[:, 2] * 0.4 > 0).astype(int)
# 1. Split training set and test set (7:3)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42
)
# 2. Train the model
model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train)
# 3. Evaluate on the test set
y_pred = model.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
print(f"Model accuracy: {accuracy:.4f}")
Output result:
模型准确率: 0.8833
Process description:

-
Advantages:Simple and fast, low computational cost.
- Disadvantages:The evaluation result is highly dependent on a single random split. If the split is unlucky, the evaluation result may not be representative. Meanwhile, since the test set is used only once, data utilization is insufficient.
2. K-Fold Cross-Validation
This is currently the most commonly used and standard cross-validation method.
Principle:Split the datasetuniformlyrandomly into K mutually exclusive subsets (called folds).
In each experiment, take turns using one subset as the test set and the remaining K-1 subsets as the training set. This process is repeated K times, ensuring each subset is used as the test set once. Finally, we obtain K evaluation scores, and take their average as the final performance estimate of the model.
Example
from sklearn.model_selection import cross_val_score, KFold
from sklearn.linear_model import LogisticRegression
# Construct runnable example data
# 100 samples, 4 features, binary classification labels
np.random.seed(42)
X = np.random.randn(100, 4)
y = (X[:, 0] + X[:, 1] * 0.5 > 0).astype(int)
# 1. Initialize the model
model = LogisticRegression(max_iter=1000)
# 2. Define the K-fold cross-validation splitter (K=5)
kfold = KFold(n_splits=5, shuffle=True, random_state=42)
# 3. Perform cross-validation
scores = cross_val_score(model, X, y, cv=kfold, scoring='accuracy')
print(f"Accuracy per fold: {scores}")
print(f"Average accuracy: {scores.mean():.4f} (+/- {scores.std() * 2:.4f})")
每次折叠的准确率: [0.9 0.95 1. 0.95 1. ] 平均准确率: 0.9600 (+/- 0.0748)
Schematic diagram of the process when K=5:

How to choose the K value?
- Common values: 5 or 10. This is an empirical trade-off.
- Smaller K value (e.g., 3): The training set is larger, but the number of evaluations is small, and the variance of the estimate may be large.
- Larger K value (e.g., 10 or 20): Evaluation is more stable (small variance), but the training set is closer to the original dataset each time, which may introduce a more optimistic estimation bias, and the computational cost increases significantly.
- Extreme case K = N (number of samples): This is leave-one-out, testing with only one sample each time. The evaluation is most unbiased, but the computational cost is extremely high, usually used only for very small datasets.
Advantages:Data is fully utilized, and evaluation results are stable and reliable.
Disadvantages:The computational cost is K times that of the hold-out method.
3. Stratified K-Fold Cross-Validation
This is an important variant of K-fold cross-validation, especially suitable forclassification problemsMediumwith imbalanced class distributiondatasets.
Problem solved:In ordinary K-fold cross-validation, random splitting may cause the proportion of a certain class in some folds to differ greatly from the original dataset. For example, if a dataset has 90% positive class and 10% negative class, randomly splitting into 5 folds may result in a fold that is entirely positive with no negative class, which makes the evaluation on that fold meaningless.
Stratified K-fold cross-validation ensures that when splitting, the proportion of samples of each class in every fold remains consistent with the overall proportion in the original dataset.
Example
# The usage is almost the same as KFold, just replace the splitter
stratified_kfold = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X, y, cv=stratified_kfold, scoring='accuracy')
For classification tasks, especially when classes are imbalanced,give priority to usingStratifiedKFold。
4. Time Series Split
Fortime series data, the order of data is crucial (tomorrow's data depends on today's and yesterday's). We cannot randomly shuffle the data; we must maintain the time order.
The principle is: the training set always consists of earlier data in time, and the test set is the data immediately following. As the number of folds increases, the training set window keeps expanding.
Example
from sklearn.model_selection import TimeSeriesSplit
# Construct runnable time series data
# 100 time points, 2 features
np.random.seed(42)
X = np.random.randn(100, 2)
# TimeSeriesSplit example
tscv = TimeSeriesSplit(n_splits=5)
for train_index, test_index in tscv.split(X):
print(f"Training set index range: {train_index)
print(f"Test set index range: {test_index)
print("---")
Output:
训练集索引范围: 0 到 19 测试集索引范围: 20 到 35 --- 训练集索引范围: 0 到 35 测试集索引范围: 36 到 51 --- 训练集索引范围: 0 到 51 测试集索引范围: 52 到 67 --- 训练集索引范围: 0 到 67 测试集索引范围: 68 到 83 --- 训练集索引范围: 0 到 83 测试集索引范围: 84 到 99 ---
Applications of Cross-Validation in Model Engineering
Cross-validation is not only an evaluation tool, but also a core part of the model optimization and engineering process.
Application 1: Model Selection and Comparison
When we need to choose among multiple candidate models (such as linear regression, decision tree, support vector machine), we cannot use the test set to select (otherwise the test set becomes part of the training process and "leaks" information). The correct approach is:
- For each candidate model, on thetraining setuse cross-validation to obtain its performance estimate.
- Compare the average scores of these cross-validation results, and select the model with the highest score.
- Finally, retrain the selected model on the entire training set, and usethe independent test setto perform a final one-time evaluation, reporting this score as the model's final performance.
Example
from sklearn.model_selection import cross_val_score, StratifiedKFold, train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
from sklearn.tree import DecisionTreeClassifier
# Construct runnable classification data
# 200 samples, 4 features, binary classification
np.random.seed(42)
X = np.random.randn(200, 4)
y = (X[:, 0] + X[:, 1] * 0.8 - X[:, 2] * 0.3 > 0).astype(int)
# Split training set / test set
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=42, stratify=y
)
# Define models
models = {
'Logistic Regression': LogisticRegression(max_iter=1000),
'SVM': SVC(),
'Decision Tree': DecisionTreeClassifier()
}
# Stratified K-fold cross-validation
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
results = {}
# Cross-validation evaluation
for name, model in models.items():
scores = cross_val_score(model, X_train, y_train, cv=cv, scoring='accuracy')
results[name] = scores.mean()
print(f"{name} average accuracy: {scores.mean():.4f}")
# Select the best model
best_model_name = max(results, key=results.get)
print(f"\nAccording to cross-validation, the best model is: {best_model_name}")
# Final training and test set evaluation
best_model = models[best_model_name]
best_model.fit(X_train, y_train)
final_score = best_model.score(X_test, y_test)
print(f"Final accuracy of the best model on the independent test set: {final_score:.4f}")
Output:
Logistic Regression 平均准确率: 0.9533 SVM 平均准确率: 0.9400 Decision Tree 平均准确率: 0.8467 根据交叉验证,最佳模型是: Logistic Regression 最佳模型在独立测试集上的最终准确率: 1.0000
Application 2: Hyperparameter Tuning
Hyperparameters are parameters that need to be set before model training (such as the number of trees in a random forestn_estimators, the penalty coefficient of SVMC). The process of finding the best hyperparameter combination is calledhyperparameter tuning, and cross-validation is its standard evaluation method.
The most commonly used method isgrid search cross-validation。
Example
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.ensemble import RandomForestClassifier
# Construct runnable classification data
# 300 samples, 5 features, binary classification
np.random.seed(42)
X = np.random.randn(300, 5)
y = (X[:, 0] * 0.6 + X[:, 1] * 0.4 - X[:, 2] > 0).astype(int)
# Split training set / test set
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=42, stratify=y
)
# 1. Parameter grid
param_grid = {
'n_estimators': [50, 100, 200],
'max_depth': [None, 10, 20],
'min_samples_split': [2, 5, 10]
}
# 2. Base model
rf = RandomForestClassifier(random_state=42)
# 3. GridSearchCV
grid_search = GridSearchCV(
estimator=rf,
param_grid=param_grid,
cv=5,
scoring='accuracy',
n_jobs=-1
)
# 4. Grid search (training set only)
grid_search.fit(X_train, y_train)
# 5. Optimal parameters and score
print(f"Best parameters: {grid_search.best_params_}")
print(f"Best cross-validation score: {grid_search.best_score_:.4f}")
# 6. Test set evaluation
best_rf_model = grid_search.best_estimator_
test_accuracy = best_rf_model.score(X_test, y_test)
print(f"Accuracy of the tuned model on the test set: {test_accuracy:.4f}")
Output:
最佳参数: {'max_depth': None, 'min_samples_split': 2, 'n_estimators': 100}
最佳交叉验证分数: 0.9067
调优后模型在测试集上的准确率: 0.9467
Key points: GridSearchCVInternally, cross-validation has already been completed. It further splits the training setX_traininto smaller "training subsets" and "validation subsets" to evaluate parameters, so what we pass inX_trainis equivalent to the entire "question bank", whileX_testis the "ultimate exam question" that has never participated in the tuning process and is used for final verification.
Practice Exercises and Summary
Hands-on Exercise
- Basic implementation: Use
sklearnthe built-in Iris dataset, respectively usingtrain_test_splitandcross_val_score(K=5) to train and evaluate aKNeighborsClassifier, and compare the scores obtained by the two evaluation methods. - Model comparison: On the same dataset, use cross-validation to compare
SVC、RandomForestClassifierandGradientBoostingClassifierthe performance of the models. - Parameter tuning: For
SVCthe model, design a parameter grid (includingCandgamma), useGridSearchCVto find the optimal parameters.
Summary of Key Points
- Purpose of cross-validation: Obtain a more robust and reliable estimate of the model's generalization ability.
- Core method:K-fold cross-validationis the gold standard; for classification problems, prioritize usingstratified K-fold cross-validation。
- Key difference:
- Training set/validation set: Used for model training andevaluation during the development process(such as selecting models, tuning parameters). Cross-validation occurs at this stage.
- Test set: Only after all development is complete, used forthe final one-time performance report. It must be strictly isolated and must never be used for any decision-making.
- Engineering role: Cross-validation is the bridge connecting model development (training, tuning) and model evaluation (testing), and is the cornerstone for ensuring model quality, preventing overfitting, and achieving reliable model selection and optimization.