Ensemble Learning

In the field of machine learning, ensemble learning is a technique that improves overall performance by combining the prediction results of multiple models.

The core idea of ensemble learning is "three cobblers with their wits combined equal Zhuge Liang," that is, combining multiple weak learners can build a strong learner.

The main goal of ensemble learning is to improve prediction accuracy and robustness by combining multiple models.

Common ensemble learning methods include:

  1. Bagging: Generate multiple training sets through bootstrap sampling, then train multiple models separately, and finally obtain the final result by voting or averaging.
  2. Boosting: Train multiple models iteratively, each model attempts to correct the errors of the previous model, and finally the result is obtained through weighted voting.
  3. Stacking: Train multiple different models, then use the outputs of these models as new features, and train a meta-model for the final prediction.

1. Bagging(Bootstrap Aggregating)

The goal of Bagging is to improve performance by reducing the variance of the model, and it is suitable for models with high variance and a tendency to overfit. It is implemented through the following steps:

  • Data set resampling: Perform multiple random samplings with replacement (bootstrap) on the training data set, each sampling yielding a subset.
  • Train multiple models: Train a base learner (usually the same type of model) on each subset.
  • Merge results: Combine the results of multiple base learners, usually by voting (for classification problems) or averaging (for regression problems).

Typical algorithms:

  • Random Forest: Random Forest is a classic implementation of Bagging. It reduces the risk of overfitting by building multiple decision trees, each tree randomly selecting features during training.

Advantages:

  • It can effectively reduce variance and improve model stability.
  • Suitable for high-variance models, such as decision trees.

Disadvantages:

  • The training process takes a long time because multiple models need to be trained.
  • The results are difficult to interpret because there is no single model.


2. Boosting

The goal of Boosting is to improve performance by reducing the bias of the model, and it is suitable for weak learners. The core idea of Boosting is to gradually adjust the weight of each model, emphasizing those samples that were misclassified by the previous round of models. Boosting is implemented through the following steps:

  • Serial training: Models are trained one after another, and each round of training is adjusted according to the errors of the previous round.
  • Weighted voting: The final prediction is the weighted sum of all weak learner predictions, where misclassified samples are assigned higher weights.
  • Combine models: The weight of each model is determined by its performance during the training process.

Typical algorithms:

  • AdaBoost(Adaptive Boosting): AdaBoost changes the weights of samples, making each subsequent classifier pay more attention to samples misclassified in the previous round.
  • Gradient Boosting Trees (GBT): GBT iteratively optimizes the objective function to gradually reduce bias.
  • XGBoost(Extreme Gradient Boosting): XGBoost is an efficient gradient boosting algorithm, widely used in data science competitions, with strong performance and optimization.
  • LightGBM(Light Gradient Boosting Machine): LightGBM is a framework based on gradient boosting trees. Compared to XGBoost, it has faster training speed and lower memory usage.

Advantages:

  • Suitable for models with large bias, and can effectively improve prediction accuracy.
  • Strong performance, and it performs excellently in many real-world applications.

Disadvantages:

  • It is sensitive to noisy data and can easily lead to overfitting.
  • The training process is slow, especially with large data sets.


3. Stacking(Stacked Generalization)

Stacking is a method that improves overall prediction accuracy by training different kinds of models and combining their predictions. Its core idea is:

  • First layer (base learners): Train multiple base learners of different types (e.g., decision trees, SVM, KNN, etc.) to predict the data.
  • Second layer (meta-learner): Use the prediction results of the first-layer learners as input, and train a meta-learner (usually logistic regression, linear regression, etc.) to make the final prediction.

Advantages:

  • Different types of base learners can be used to capture different patterns in the data.
  • In theory, it can combine the advantages of multiple models to achieve stronger prediction capability.

Disadvantages:

  • The training process is complex, requiring training of multiple models, and the way models are combined also needs careful design.
  • It is more complex than other ensemble methods such as Bagging and Boosting, and is prone to overfitting.


Example demonstration

Example

from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

# Load the data set
iris = load_iris()
X, y = iris.data, iris.target

# Split the data set into training and test sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

# Create a random forest classifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)

# Train the model
rf.fit(X_train, y_train)

# Predict
y_pred = rf.predict(X_test)

# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print(f"Random forest accuracy: {accuracy:.2f}")

The output result is as follows:

随机森林的准确率: 1.00

Code explanation:

  • Load the data set: We use theload_iris()function to load the classic Iris data set.
  • Split the data set into training and test sets: Use thetrain_test_split()function to divide the data set into training and test sets.
  • Create a random forest classifier: Use theRandomForestClassifierclass to create a random forest classifier,n_estimators=100meaning 100 decision trees are used.
  • Train the model: Use thefit()method to train the model.
  • Predict: Use thepredict()method to make predictions on the test set.
  • Calculate accuracy: Use theaccuracy_score()function to calculate the accuracy of the model.

Boosting:AdaBoost

Algorithm principle:The core idea of Boosting is to train multiple models iteratively, with each model attempting to correct the errors of the previous model. AdaBoost (Adaptive Boosting) is one of the most classic Boosting algorithms.

Example

from sklearn.ensemble import AdaBoostClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
from sklearn.tree import DecisionTreeClassifier

# Load the data set
iris = load_iris()
X, y = iris.data, iris.target

# Split the data set into training and test sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

# Use the default weak learner (decision tree) and specify the SAMME algorithm
ada = AdaBoostClassifier(DecisionTreeClassifier(max_depth=1),
                         n_estimators=50,
                         random_state=42,
                         algorithm='SAMME')

# Train the model
ada.fit(X_train, y_train)

# Predict
y_pred = ada.predict(X_test)

# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print(f"AdaBoost accuracy: {accuracy:.2f}")

The output result is as follows:

AdaBoost的准确率: 1.00

Code explanation:

  1. Load the data set: Use theload_iris()function to load the Iris data set, containing feature dataXand label data.y。
  2. Split the data set into training and test sets: Use thetrain_test_split()function to split the data set into training and test sets, with the test set accounting for 30% and the training set accounting for 70%.
  3. Create a decision tree classifier: UseDecisionTreeClassifier(max_depth=1)to create a decision tree classifier with depth 1, as the base learner of AdaBoost.
  4. Create an AdaBoost classifier: Use theAdaBoostClassifier()class to create an AdaBoost classifier,n_estimators=50meaning 50 weak learners are used,algorithm='SAMME'and the SAMME algorithm is specified.
  5. Train the model: Use thefit()method to train the AdaBoost model on the training data.
  6. Predict: Use thepredict()method to make predictions on the test set, generating predicted labels.y_pred。
  7. Calculate accuracy: Use theaccuracy_score()function to calculate and output the prediction accuracy of the model.

Stacking: Model Stacking

Algorithm principle:The core idea of Stacking is to train multiple different models, then use the outputs of these models as new features, and train a meta-model to make the final prediction.

Example

from sklearn.ensemble import StackingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

# Load the data set
iris = load_iris()
X, y = iris.data, iris.target

# Split the data set into training and test sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

# Define base learners
estimators = [
    ('dt', DecisionTreeClassifier(max_depth=1)),
    ('svc', SVC(kernel='linear', probability=True))
]

# Create a Stacking classifier
stacking = StackingClassifier(estimators=estimators, final_estimator=LogisticRegression())

# Train the model
stacking.fit(X_train, y_train)

# Predict
y_pred = stacking.predict(X_test)

# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print(f"Stacking accuracy: {accuracy:.2f}")

The output result is as follows:

tacking的准确率: 1.00

Code explanation:

  1. Load the data set: Likewise, use theload_iris()The function loads the Iris dataset.
  2. Split the training set and test set: Usetrain_test_split()The function divides the dataset into training and test sets.
  3. Define the base learner: UseDecisionTreeClassifierandSVCas the base learner.
  4. Create a Stacking classifier: UseStackingClassifierclass to create a Stacking classifier,final_estimator=LogisticRegression()indicating logistic regression is used as the meta-model.
  5. Train the model: Usefit()method to train the model.
  6. Predict: Usepredict()method to make predictions on the test set.
  7. Calculate accuracy: Useaccuracy_score()function to calculate the model's accuracy.
Other extensions