Sklearn Custom Models and Features

Custom models can extend scikit-learn's functionality to meet the needs of specific application scenarios.

In scikit-learn, in addition to using ready-made models and preprocessing functions, users can also create custom models, transformers, and features according to their own needs.

Implementing custom models and features usually involves inheriting scikit-learn's base classes, such as BaseEstimator and TransformerMixin, and then implementing specific fit and predict methods.

This chapter will elaborate on the following aspects:

  1. Custom Transformer
  2. Custom Estimator
  3. Custom Pipeline Steps
  4. How to use inheritanceBaseEstimatorandTransformerMixinto implement custom algorithms

1. Custom Transformer

A transformer is a component used to transform data, such as standardization, feature selection, etc. Custom transformers can inheritTransformerMixin, and implementfitandtransformmethods.

Steps for customizing a transformer:

  • fitmethod: typically used to learn data attributes (such as mean, variance, feature selection criteria, etc.).fitThe method returns the transformer itself to allow chained calls.
  • transformmethod: applies the learned attributes to transform or process the data.

Custom transformer example: custom standardization transformer:

Suppose we want to implement a custom standardization transformer. Standardization scales each feature of the data to have a mean of 0 and a variance of 1.

Example

import numpy as np
from sklearn.base import BaseEstimator, TransformerMixin

class CustomScaler(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        """
Compute the mean and standard deviation of each feature
        """

        self.mean_ = np.mean(X, axis=0)
        self.std_ = np.std(X, axis=0)
        return self  # Return the object itself

    def transform(self, X):
        """
Standardize the data
        """

        return (X - self.mean_) / self.std_

# Test the custom transformer
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

# Load data
data = load_iris()
X, y = data.data, data.target
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Create a custom transformer object
scaler = CustomScaler()

# Use custom standardization
scaler.fit(X_train)
X_train_scaled = scaler.transform(X_train)
X_test_scaled = scaler.transform(X_test)

print("Scaled training data:\n", X_train_scaled)

The output is as follows:

Scaled training data:
 [[-1.47393679  1.20365799 -1.56253475 -1.31260282]
 [-0.13307079  2.99237573 -1.27600637 -1.04563275]
 [ 1.08589829  0.08570939  0.38585821  0.28921757]
 [-1.23014297  0.75647855 -1.2187007  -1.31260282]
 [-1.7177306   0.30929911 -1.39061772 -1.31260282]
....

Explanation:

The CustomScaler class implements a standardization process similar to StandardScaler in scikit-learn.

  • fitThe method computes the mean and standard deviation of the training data and saves these values.
  • transformThe method uses thefitmean and standard deviation calculated in the fit method to transform the data.

2. Custom Estimator

An estimator refers to the model itself, such as regressors, classifiers, etc.

Custom estimators need to inherit the BaseEstimator class and implement the fit and predict methods.

Steps for a custom estimator:

  • fitmethod: used to train the model and compute required parameters (such as weights, biases, etc.).
  • predictmethod: predicts input data based on the trained parameters.

Custom Estimator Example: Simple Classifier

Suppose we want to implement a very simple classifier: use the mean of each feature as a threshold; if a feature value exceeds the mean, predict class 1; otherwise, predict class 0.

Example

from sklearn.base import BaseEstimator
import numpy as np

class SimpleClassifier(BaseEstimator):
    def fit(self, X, y):
        """
Train the model: compute the mean of each feature
        """

        self.mean_ = np.mean(X, axis=0)
        return self  # Return the object itself

    def predict(self, X):
        """
Classify based on the mean: if the feature value is greater than the mean, predict 1; otherwise, predict 0
        """

        return (X > self.mean_).astype(int)

# Test the custom classifier
X_train = np.array([[1.5, 2.5], [2.0, 3.0], [3.5, 4.5], [4.0, 5.0]])
y_train = np.array([0, 0, 1, 1])

# Create a custom classifier object
classifier = SimpleClassifier()

# Train the model
classifier.fit(X_train, y_train)

# Make predictions
X_test = np.array([[2.5, 3.5], [1.0, 2.0]])
y_pred = classifier.predict(X_test)

print("Predictions:", y_pred)

The output is as follows:

Predictions: [[0 0]
 [0 0]]

Explanation:

  • fitThe method computes the mean of the training data and stores it inself.mean_the attribute.
  • predictThe method makes classification predictions by comparing the test data with the mean.

This type of custom classifier is not commonly used in practice, but it demonstrates how to use BaseEstimator to create a simple model.


3. Custom Pipeline Steps

scikit-learn allows custom transformers and estimators to be used as pipeline steps. This makes it possible to integrate data preprocessing and model training into a single workflow. Custom models and features can work with pipelines just like built-in transformers and estimators.

Using Custom Transformers and Estimators in a Pipeline

Example

import numpy as np
from sklearn.base import BaseEstimator, TransformerMixin
from sklearn.pipeline import Pipeline


class CustomScaler(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        """
Compute the mean and standard deviation of each feature
        """

        self.mean_ = np.mean(X, axis=0)
        self.std_ = np.std(X, axis=0)
        return self  # Return the object itself

    def transform(self, X):
        """
Standardize the data
        """

        return (X - self.mean_) / self.std_
   

class SimpleClassifier(BaseEstimator):
    def fit(self, X, y):
        """
Train the model: compute the mean of each feature
        """

        self.mean_ = np.mean(X, axis=0)
        return self  # Return the object itself

    def predict(self, X):
        """
Classify based on the mean: if the feature value is greater than the mean, predict 1; otherwise, predict 0
        """

        return (X > self.mean_).astype(int)

# Test the custom classifier
X_train = np.array([[1.5, 2.5], [2.0, 3.0], [3.5, 4.5], [4.0, 5.0]])
y_train = np.array([0, 0, 1, 1])

# Create a pipeline containing custom standardization and classifier
pipeline = Pipeline([
    ('scaler', CustomScaler()),  # Custom standardization
    ('classifier', SimpleClassifier())  # Custom classifier
])

# Train the pipeline
pipeline.fit(X_train, y_train)

X_test = np.array([[2.5, 3.5], [1.0, 2.0]])
# Predict
y_pred = pipeline.predict(X_test)
print("Predictions:", y_pred)

The output is as follows:

Predictions: [[0 0]
 [0 0]]

Explanation:

We use CustomScaler and SimpleClassifier as pipeline steps, leveraging the pipeline to automatically perform data preprocessing and model training.

In this way, the entire pipeline workflow can be completed with a single fit and predict method, maintaining efficiency and simplicity.


How to Implement Custom Algorithms by Inheriting BaseEstimator and TransformerMixin

  • BaseEstimator: is the base class for all scikit-learn estimators. It providesget_params()andset_params()methods, which allows custom estimators to work togetherGridSearchCVwith other tools.
  • TransformerMixin: is the base class for all scikit-learn transformers. It providesfit_transform()methods, which allows transformers to be used togetherPipelinewith the pipeline.

By inheriting these two base classes, we can very conveniently create our own custom models and features, and use scikit-learn's tools (such as GridSearchCV and Pipeline) for tuning and evaluation.

Complete custom transformer and estimator example:

Example

from sklearn.base import BaseEstimator, TransformerMixin
import numpy as np

class CustomEstimator(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        # Simulate a simple "model": compute the mean of each feature
        self.mean_ = np.mean(X, axis=0)
        return self

    def transform(self, X):
        # Standardize the data based on the mean
        return X - self.mean_

    def predict(self, X):
        # Simple prediction method: if the feature value is greater than the mean, predict 1; otherwise, predict 0
        return (X > self.mean_).astype(int)

# Use custom estimator and transformer
from sklearn.pipeline import Pipeline
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

# Load data
data = load_iris()
X, y = data.data, data.target
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Create pipeline
pipeline = Pipeline([
    ('custom', CustomEstimator())
])

# Train the model
pipeline.fit(X_train, y_train)

# Predict
y_pred = pipeline.predict(X_test)
print("Predictions:", y_pred)

In the above example, we created a custom estimator, CustomEstimator, which not only performs standardization (transform) but also performs prediction (predict). Then, we used it as part of a pipeline for training and prediction.

The output is as follows:

Predictions: [[1 0 1 1]
 [0 1 0 0]
 [1 0 1 1]
 [1 0 1 1]
 [1 0 1 1]
 [0 1 0 0]
 [0 0 0 1]
 [1 1 1 1]
 [1 0 1 1]
 [0 0 1 1]
...
Other Extensions