Sklearn Data Preprocessing

Data preprocessing is a critical step in machine learning projects; it directly affects the model's training effectiveness and final performance.

When building machine learning models, data preprocessing is a crucial step; it helps us clean and transform raw data to provide the best input for machine learning models.

Data preprocessing involves multiple steps, including handling missing values, data transformation, standardization, encoding, etc.

Appropriate preprocessing not only improves model accuracy but also helps the model generalize better.


1. Handling Missing Values

Missing values refer to the absence of values for certain features in the dataset.

Machine learning algorithms usually cannot directly handle missing values, so we need to process them.

Check Missing Values

First, check whether the dataset contains missing values.

Typically, we can use pandas to inspect missing values in the dataset:

Example

import pandas as pd

# Assume we have a DataFrame df
print(df.isnull().sum())  # View the number of missing values in each column

Fill Missing Values

The most common method for handling missing values is imputation (filling).

Common imputation strategies include:

  • Mean imputation: Suitable for numerical data.
  • Median imputation: For datasets with outliers, using the median may be more effective.
  • Mode imputation (most frequent value): Suitable for categorical data.

In scikit-learn, SimpleImputer can easily implement missing value imputation:

Example

from sklearn.impute import SimpleImputer

# For numerical data, use mean imputation
imputer = SimpleImputer(strategy='mean')  # Options: 'mean', 'median', 'most_frequent'
df_imputed = imputer.fit_transform(df)  # Impute missing values

Delete Missing Values

If the number of missing values is small, and deleting them does not significantly affect the analysis results, another option is to directly delete the missing values.

df_cleaned = df.dropna()  # 删除包含缺失值的行
For more details, refer to:Pandas Data Cleaning

2. Data Scaling

Machine learning algorithms are sensitive to the scale of data, so the data needs to be scaled so that features have the same scale.

Common scaling methods include:

  • Standardization: Transform the data to have a mean of 0 and a standard deviation of 1. Suitable for most machine learning algorithms.
  • Normalization: Scale the data to a specified range (usually [0, 1]).

Standardization

Standardization can be implemented with StandardScaler, which transforms each feature to have zero mean and unit variance:

Example

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)  # Standardize X

Normalization

Normalization scales each feature to a specified range (usually [0, 1]).

MinMaxScaler is used to normalize data:

Example

from sklearn.preprocessing import MinMaxScaler

scaler = MinMaxScaler()
X_normalized = scaler.fit_transform(X)  # Normalize X

Why Do We Need Standardization and Normalization?

  • Standardization: It is very important for distance-based metrics (such as K-Nearest Neighbors, Support Vector Machines, etc.), because inconsistent feature scales may cause certain features to have too large an influence on the model. Standardization ensures that each feature contributes equally to the model.

  • Normalization: Some algorithms (such as neural networks, gradient descent optimization algorithms, etc.) are very sensitive to the range of input data. Normalization helps accelerate convergence.


3. Categorical Variable Encoding

Machine learning models usually cannot directly handle string-type categorical variables, so categorical variables need to be converted into numerical data.

Common encoding methods include:

Label Encoding

Label encoding maps each category to a unique integer.

Suitable for cases where there is an ordinal relationship between categories (e.g., low, medium, high).

Example

from sklearn.preprocessing import LabelEncoder

# Assume we have a categorical variable y
label_encoder = LabelEncoder()
y_encoded = label_encoder.fit_transform(y)  # Convert the categorical variable to integers

One-Hot Encoding

One-hot encoding converts each category into a binary vector, suitable for cases where there is no ordinal relationship between categories (e.g., colors, countries, etc.).

OneHotEncoder can convert categorical variables into one-hot encoding.

Example

from sklearn.preprocessing import OneHotEncoder

# Assume we have a categorical variable X
encoder = OneHotEncoder(sparse=False)  # sparse=False returns a dense matrix
X_encoded = encoder.fit_transform(X)  # Convert the categorical variable to one-hot encoding

In pandas, the get_dummies() function can also be used for one-hot encoding:

X_encoded = pd.get_dummies(X)

4. Feature Selection

Feature selection improves model performance by selecting the most important features and reduces computational cost.

Common feature selection methods include:

Model-Based Feature Selection

Use some machine learning models (such asdecision treesorrandom forests) to evaluate feature importance, thereby performing feature selection.

Example

from sklearn.ensemble import RandomForestClassifier

# Train a random forest model
clf = RandomForestClassifier()
clf.fit(X_train, y_train)

# Get feature importance
importances = clf.feature_importances_
print(importances)

Recursive Feature Elimination (RFE)

RFE is a method that recursively deletes the least important features step by step to select the optimal features.

RFE can help us automatically select important features.

Example

from sklearn.feature_selection import RFE

# Use a linear model for recursive feature elimination
rfe = RFE(clf, n_features_to_select=3)  # Retain the 3 most important features
X_rfe = rfe.fit_transform(X_train, y_train)

5. Feature Engineering

Feature engineering aims to improve model performance by processing, combining, or constructing new features from existing ones.

Common methods of feature engineering include:

  • Feature Combination: Combine two or more features into a new feature.
  • Feature Transformation: Apply transformations such as logarithmic or square root transformations to features to address non-linearity in data.
  • Feature Creation: Create new features from existing data, such as extracting date, month, weekday, etc., from a timestamp.

Example

# For example, combine two numerical features into a new feature
df['new_feature'] = df['feature1'] * df['feature2']

6. Feature Extraction

Feature extraction aims to extract new, more expressive features from original features.

Common feature extraction methods include:Principal Component Analysis (PCA)andLinear Discriminant Analysis (LDA)。

Principal Component Analysis (PCA)

PCA is a commonly used dimensionality reduction technique that maps data from a high-dimensional space to a lower-dimensional space via linear transformation, making the new features (principal components) preserve as much variance as possible.

PCA is especially suitable when there are too many features, as it can effectively reduce computational complexity.

Example

from sklearn.decomposition import PCA

# Assume X is the feature matrix
pca = PCA(n_components=2)  # Reduce to 2 principal components
X_pca = pca.fit_transform(X)

PCA is mainly used in two scenarios:

  • Dimensionality Reduction: When there are too many features, using PCA for dimensionality reduction can reduce computational cost while retaining the main information of the data.
  • Visualization: Map high-dimensional data to 2D or 3D space to help us visualize the data structure.

Linear Discriminant Analysis (LDA)

LDA is a supervised learning dimensionality reduction method that aims to find a linear combination that maximizes the distance between different classes and minimizes the distance within classes.

LDA is commonly used in classification tasks.

Example

from sklearn.discriminant_analysis import LinearDiscriminantAnalysis

# Assume X is the feature matrix and y is the target variable
lda = LinearDiscriminantAnalysis(n_components=2)  # Reduce dimensionality to 2 linear discriminant components
X_lda = lda.fit_transform(X, y)

7. Handling Imbalanced Data

In classification problems, if the number of samples across classes in the dataset varies greatly, the model may become biased toward predicting the majority class, thereby affecting model performance.

Common handling methods include:

Over-sampling

Balance the dataset by increasing the number of samples in the minority class.

A common method is to use the SMOTE (Synthetic Minority Over-sampling Technique) algorithm.

Example

from imblearn.over_sampling import SMOTE

smote = SMOTE()
X_resampled, y_resampled = smote.fit_resample(X_train, y_train)

Under-sampling

Balance the dataset by reducing the number of samples in the majority class.

Example

from imblearn.under_sampling import RandomUnderSampler

undersampler = RandomUnderSampler()
X_resampled, y_resampled = undersampler.fit_resample(X_train, y_train)
Other Extensions