Sklearn Basic Concepts

When using Sklearn for machine learning, you need to understand some basic concepts.

Sklearn provides a unified and concise API to implement various machine learning algorithms and workflows, helping us quickly accomplish various machine learning tasks.

Next, we will elaborate on the following concepts:Data representation, model types, preprocessing methods, evaluation metrics, model tuningetc.


1. Data Representation: Datasets and Features

The dataset is one of the most basic concepts in Sklearn.

The core task of machine learning is to learn patterns from data, and the way data is represented is crucial.

Dataset

In scikit-learn, data is usually represented by two main objects:Feature matrixandTarget vector。

Feature Matrix:Each row represents a data sample, and each column represents a feature (i.e., an input variable). It is a two-dimensional array or matrix, usually stored as a NumPy array or pandas DataFrame.

Suppose we have 3 samples, each with 2 features.

Example

import numpy as np
X = np.array([[1.0, 2.0], [2.0, 3.0], [3.0, 4.0]])

Target Vector:It represents the target of each sample (i.e., the output label), usually a one-dimensional array.

For example, in a classification task, the target is the class label of each sample.

The corresponding target vector:

y = np.array([0, 1, 0])  # 0 类别和 1 类别

Features and Labels

  • Features: are the input variables in the dataset used to train the model. In the example above,Xis the feature matrix, containing all input variables.

  • Labels: are the target outputs of the machine learning model. In supervised learning, labels are the results we want the model to predict. In the example above,yis the label or target vector, containing the class of each sample.

Dataset Splitting

In practical applications, it is usually necessary to split the dataset into a training set and a test set.

scikit-learn provides a convenient functiontrain_test_split()to achieve this:

Example

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
  • The above code callstrain_test_splitthe function, and assigns the result to four variables:X_train、X_test、y_trainandy_test。
  • Xandyare passed totrain_test_splitthe function parameters, which respectively represent the feature dataset and the target variable (labels). UsuallyXis a two-dimensional array,yis a one-dimensional array.
  • test_size=0.3The parameter specifies that the size of the test set should be 30% of the original dataset. This means that 70% of the data will be used as the training set, and the remaining 30% will be used as the test set.
  • random_state=42The parameter is a random seed, used to ensure that the same result is obtained each time the dataset is split. This is very useful in experiments and model validation, because it ensures the reproducibility of results.

2. Models and Algorithms

Supervised and Unsupervised Learning

In scikit-learn, machine learning models are roughly divided into two major categories:Supervised learningandUnsupervised learning。

Supervised Learning:In supervised learning, the model learns from labeled data during training; these labels are the results we want the model to predict.

Common supervised learning tasks include classification and regression.

  • Classification: Assign data points to predefined categories. For example, determining whether an email is spam or not spam.
  • Regression: Predict continuous value outputs. For example, predicting house prices, temperatures, etc.

Using a decision tree for a classification task:

Example

from sklearn.tree import DecisionTreeClassifier
clf = DecisionTreeClassifier()
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)

Unsupervised Learning:Unsupervised learning refers to learning without labeled data; the model learns only from the characteristics of the input data itself.

Common unsupervised learning tasks include clustering and dimensionality reduction.

  • Clustering: Group data so that data within the same group are similar. Common clustering algorithms includeK-Means、DBSCANetc.
  • Dimensionality Reduction: Reduce the number of features in the data, often used for data compression or visualization. Common dimensionality reduction methods includePCA(Principal Component Analysis) andt-SNE(t-distributed Stochastic Neighbor Embedding), etc.

Using K-Means clustering:

Example

from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3)
kmeans.fit(X_train)
y_pred = kmeans.predict(X_test)

Preprocessing and Feature Engineering

Before using scikit-learn for machine learning, it is usually necessary to preprocess the data. This includes the following common tasks:

1. Standardization:Unify the scale of features so that each feature has zero mean and unit variance.

Example

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

2. Normalization:Scale the values of features to a fixed range (usually 0 to 1).

Example

from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
X_normalized = scaler.fit_transform(X)

3. Categorical Variable Encoding:Convert categorical data into numerical data (e.g., one-hot encoding).


3. Model Evaluation and Validation

After training a machine learning model, its performance must be evaluated to ensure its generalization ability.

scikit-learn provides some tools for evaluating model performance.

Cross-validation

Cross-validation is a common model evaluation method, especially when the amount of data is limited.

By dividing the data into multiple subsets, using one subset as the validation set each time and the rest as the training set, the model is trained and evaluated repeatedly multiple times, and finally the average performance of the model is calculated.

Example

from sklearn.model_selection import cross_val_score
scores = cross_val_score(clf, X, y, cv=5)  # 5-fold cross-validation
print("Cross-validation scores:", scores)

Common Evaluation Metrics

Evaluation metrics for classification tasks:

  • Accuracy: The proportion of correctly predicted samples among all samples.
  • Precision: The proportion of actual positive classes among positive class predictions.
  • Recall: The proportion of correctly predicted samples among actual positive classes.
  • F1 Score: The harmonic mean of precision and recall.

Example

from sklearn.metrics import accuracy_score, classification_report
print("Accuracy:", accuracy_score(y_test, y_pred))
print("Classification Report:\n", classification_report(y_test, y_pred))

Evaluation metrics for regression tasks:

  • Mean Squared Error (MSE): The average of the squared differences between predicted values and true values.
  • Coefficient of Determination (R²): Measures the model's ability to explain the variability of the data.

Example

from sklearn.metrics import mean_squared_error, r2_score
print("MSE:", mean_squared_error(y_test, y_pred))
print("R²:", r2_score(y_test, y_pred))

4. Model Selection and Tuning

Grid Search

Grid searchis a commonly used hyperparameter tuning method that finds the best hyperparameter combination by traversing all possible parameter combinations.

Example

from sklearn.model_selection import GridSearchCV

param_grid = {'max_depth': [3, 5, 7], 'min_samples_split': [2, 5, 10]}
grid_search = GridSearchCV(DecisionTreeClassifier(), param_grid, cv=5)
grid_search.fit(X_train, y_train)
print("Best parameters:", grid_search.best_params_)

Random Search

Random searchis a method that searches for optimal hyperparameters by randomly selecting combinations of hyperparameters. It is more efficient than grid search.

Example

from sklearn.model_selection import RandomizedSearchCV
from scipy.stats import randint

param_dist = {'max_depth': [3, 5, 7], 'min_samples_split': randint(2, 10)}
random_search = RandomizedSearchCV(DecisionTreeClassifier(), param_dist, n_iter=10, cv=5)
random_search.fit(X_train, y_train)
print("Best parameters:", random_search.best_params_)
Other Extensions