Sklearn Basic Concepts
When using Sklearn for machine learning, you need to understand some basic concepts.
Sklearn provides a unified and concise API to implement various machine learning algorithms and workflows, helping us quickly accomplish various machine learning tasks.
Next, we will elaborate on the following concepts:Data representation, model types, preprocessing methods, evaluation metrics, model tuningetc.
1. Data Representation: Datasets and Features
The dataset is one of the most basic concepts in Sklearn.
The core task of machine learning is to learn patterns from data, and the way data is represented is crucial.
Dataset
In scikit-learn, data is usually represented by two main objects:Feature matrixandTarget vector。
Feature Matrix:Each row represents a data sample, and each column represents a feature (i.e., an input variable). It is a two-dimensional array or matrix, usually stored as a NumPy array or pandas DataFrame.
Suppose we have 3 samples, each with 2 features.
Example
X = np.array([[1.0, 2.0], [2.0, 3.0], [3.0, 4.0]])
Target Vector:It represents the target of each sample (i.e., the output label), usually a one-dimensional array.
For example, in a classification task, the target is the class label of each sample.
The corresponding target vector:
y = np.array([0, 1, 0]) # 0 类别和 1 类别
Features and Labels
Features: are the input variables in the dataset used to train the model. In the example above,
Xis the feature matrix, containing all input variables.Labels: are the target outputs of the machine learning model. In supervised learning, labels are the results we want the model to predict. In the example above,
yis the label or target vector, containing the class of each sample.
Dataset Splitting
In practical applications, it is usually necessary to split the dataset into a training set and a test set.
scikit-learn provides a convenient functiontrain_test_split()to achieve this:
Example
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
- The above code calls
train_test_splitthe function, and assigns the result to four variables:X_train、X_test、y_trainandy_test。 Xandyare passed totrain_test_splitthe function parameters, which respectively represent the feature dataset and the target variable (labels). UsuallyXis a two-dimensional array,yis a one-dimensional array.test_size=0.3The parameter specifies that the size of the test set should be 30% of the original dataset. This means that 70% of the data will be used as the training set, and the remaining 30% will be used as the test set.random_state=42The parameter is a random seed, used to ensure that the same result is obtained each time the dataset is split. This is very useful in experiments and model validation, because it ensures the reproducibility of results.
2. Models and Algorithms
Supervised and Unsupervised Learning
In scikit-learn, machine learning models are roughly divided into two major categories:Supervised learningandUnsupervised learning。
Supervised Learning:In supervised learning, the model learns from labeled data during training; these labels are the results we want the model to predict.
Common supervised learning tasks include classification and regression.
- Classification: Assign data points to predefined categories. For example, determining whether an email is spam or not spam.
- Regression: Predict continuous value outputs. For example, predicting house prices, temperatures, etc.
Using a decision tree for a classification task:
Example
clf = DecisionTreeClassifier()
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
Unsupervised Learning:Unsupervised learning refers to learning without labeled data; the model learns only from the characteristics of the input data itself.
Common unsupervised learning tasks include clustering and dimensionality reduction.
- Clustering: Group data so that data within the same group are similar. Common clustering algorithms includeK-Means、DBSCANetc.
- Dimensionality Reduction: Reduce the number of features in the data, often used for data compression or visualization. Common dimensionality reduction methods includePCA(Principal Component Analysis) andt-SNE(t-distributed Stochastic Neighbor Embedding), etc.
Using K-Means clustering:
Example
kmeans = KMeans(n_clusters=3)
kmeans.fit(X_train)
y_pred = kmeans.predict(X_test)
Preprocessing and Feature Engineering
Before using scikit-learn for machine learning, it is usually necessary to preprocess the data. This includes the following common tasks:
1. Standardization:Unify the scale of features so that each feature has zero mean and unit variance.
Example
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
2. Normalization:Scale the values of features to a fixed range (usually 0 to 1).
Example
scaler = MinMaxScaler()
X_normalized = scaler.fit_transform(X)
3. Categorical Variable Encoding:Convert categorical data into numerical data (e.g., one-hot encoding).
3. Model Evaluation and Validation
After training a machine learning model, its performance must be evaluated to ensure its generalization ability.
scikit-learn provides some tools for evaluating model performance.
Cross-validation
Cross-validation is a common model evaluation method, especially when the amount of data is limited.
By dividing the data into multiple subsets, using one subset as the validation set each time and the rest as the training set, the model is trained and evaluated repeatedly multiple times, and finally the average performance of the model is calculated.
Example
scores = cross_val_score(clf, X, y, cv=5) # 5-fold cross-validation
print("Cross-validation scores:", scores)
Common Evaluation Metrics
Evaluation metrics for classification tasks:
- Accuracy: The proportion of correctly predicted samples among all samples.
- Precision: The proportion of actual positive classes among positive class predictions.
- Recall: The proportion of correctly predicted samples among actual positive classes.
- F1 Score: The harmonic mean of precision and recall.
Example
print("Accuracy:", accuracy_score(y_test, y_pred))
print("Classification Report:\n", classification_report(y_test, y_pred))
Evaluation metrics for regression tasks:
- Mean Squared Error (MSE): The average of the squared differences between predicted values and true values.
- Coefficient of Determination (R²): Measures the model's ability to explain the variability of the data.
Example
print("MSE:", mean_squared_error(y_test, y_pred))
print("R²:", r2_score(y_test, y_pred))
4. Model Selection and Tuning
Grid Search
Grid searchis a commonly used hyperparameter tuning method that finds the best hyperparameter combination by traversing all possible parameter combinations.
Example
param_grid = {'max_depth': [3, 5, 7], 'min_samples_split': [2, 5, 10]}
grid_search = GridSearchCV(DecisionTreeClassifier(), param_grid, cv=5)
grid_search.fit(X_train, y_train)
print("Best parameters:", grid_search.best_params_)
Random Search
Random searchis a method that searches for optimal hyperparameters by randomly selecting combinations of hyperparameters. It is more efficient than grid search.
Example
from scipy.stats import randint
param_dist = {'max_depth': [3, 5, 7], 'min_samples_split': randint(2, 10)}
random_search = RandomizedSearchCV(DecisionTreeClassifier(), param_dist, n_iter=10, cv=5)
random_search.fit(X_train, y_train)
print("Best parameters:", random_search.best_params_)