Sklearn House Price Prediction

Next, we use sklearn to predict Beijing housing prices.

We will proceed step by step from data loading, exploratory analysis, and feature engineering to model training and optimization, demonstrating how to use sklearn's tools and libraries to complete the prediction task.

Overview:

1. Data Generation and Inspection

  • A dictionary is used to create simulated data, and then it is converted intopandas DataFrame。
  • The data contains the house area, number of rooms, floor, year built, location (categorical variable), and house price (target variable).
  • Throughdf.head()anddf.describe()to check the basic structure and statistical information of the data.

2. Data Preprocessing

  • Feature selection: Features related to house price (area, number of rooms, floor, year built, location) are selected from the original data.
  • Data splitting: Usetrain_test_splitto split the dataset into 80% training set and 20% test set.
  • Numerical feature standardization: UseStandardScalerto standardize numerical features.
  • Categorical feature encoding: UseOneHotEncoderfor categorical features (location) perform One-Hot encoding.
  • ColumnTransformerIntegrate the processing of numerical and categorical features into one step.

3. Model Training

  • UsePipelineto combine data preprocessing with model training steps, ensuring the entire process is pipelined.
  • Use the linear regression model (LinearRegression) for training.

4. Model Evaluation

  • Bymean_squared_errorcalculating the mean squared error (MSE), and byr2_scorecalculating the coefficient of determination (R²).
  • Output the evaluation results to view the model's prediction accuracy.

5. Model Optimization

  • UseGridSearchCVto tune the hyperparameters of linear regression. The main adjustment isfit_intercept(whether to fit the intercept).
  • Find the best hyperparameters through grid search, and use the best model to predict on the test set.
  • Recalculate the evaluation metrics (MSE and R²) of the optimized model.

1. Data Generation and Inspection

We first construct a simulated DataFrame containing some common house price prediction features, such as house area, number of rooms, floor, year built, and geographic location (categorical variable).

Example

import pandas as pd
import numpy as np

# Simulated data: house area (square meters), number of rooms, floor, year built, location (categorical variable)
data = {
    'area': [70, 85, 100, 120, 60, 150, 200, 80, 95, 110],
    'rooms': [2, 3, 3, 4, 2, 5, 6, 3, 3, 4],
    'floor': [5, 2, 8, 10, 3, 15, 18, 7, 9, 11],
    'year_built': [2005, 2010, 2012, 2015, 2000, 2018, 2020, 2008, 2011, 2016],
    'location': ['Chaoyang', 'Haidian', 'Chaoyang', 'Dongcheng', 'Fengtai', 'Haidian', 'Chaoyang', 'Fengtai', 'Dongcheng', 'Haidian'],
    'price': [5000000, 6000000, 6500000, 7000000, 4500000, 10000000, 12000000, 5500000, 6200000, 7500000]  # House price (target variable)
}

# Create DataFrame
df = pd.DataFrame(data)

# View data
print("Data preview:")
print(df.head())

Output:

数据预览:
   area  rooms  floor  year_built location    price
0    70      2      5        2005   Chaoyang  5000000
1    85      3      2        2010   Haidian  6000000
2   100      3      8        2012   Chaoyang  6500000
3   120      4     10        2015  Dongcheng  7000000
4    60      2      3        2000   Fengtai  4500000

2. Data Preprocessing

Data preprocessing usually includes feature selection, feature transformation, missing value handling, data standardization, and so on. We will standardize numerical features and one-hot encode categorical features.

Example

from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline

# Feature selection
X = df[['area', 'rooms', 'floor', 'year_built', 'location']]  # Features
y = df['price']  # Target variable

# Split training set and test set
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Build preprocessing steps
numeric_features = ['area', 'rooms', 'floor', 'year_built']
categorical_features = ['location']

numeric_transformer = Pipeline(steps=[
    ('scaler', StandardScaler())  # Standardize numerical features
])

categorical_transformer = Pipeline(steps=[
    ('onehot', OneHotEncoder(handle_unknown='ignore'))  # Handle new categories in the test set
])

# Combine into ColumnTransformer
preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, numeric_features),
        ('cat', categorical_transformer, categorical_features)
    ]
)

# View the structure after data preprocessing
X_train_transformed = preprocessor.fit_transform(X_train)
print("Preprocessed training data:")
print(X_train_transformed)

The output result is:

tianqixin@Mac-mini example-test % python3 test.py
预处理后的训练数据:
[[ 0.89826776  1.0440738   1.14636101  0.96800387  0.          0.
   0.          1.        ]
 [-0.95622052 -1.23390539 -0.98640366 -1.04544418  1.          0.
   0.          0.        ]
 [-0.72440948 -0.474579   -0.55985073 -0.58080232  0.          0.
   1.          0.        ]
 [-0.26078741 -0.474579   -0.34657426  0.03872015  1.          0.
   0.          0.        ]
 [-0.02897638  0.2847474   0.29325514  0.65824263  0.          0.
   0.          1.        ]
 [-1.18803155 -1.23390539 -1.4129566  -1.81984727  0.          0.
   1.          0.        ]
 [ 0.20283466  0.2847474   0.07997868  0.50336201  0.          1.
   0.          0.        ]
 [ 2.05732294  1.80340019  1.78619041  1.2777651   1.          0.
   0.          0.        ]]

3. Model Building

Next, we use a linear regression model to predict house prices. We use a Pipeline to integrate the preprocessing and model training steps.

Example

from sklearn.linear_model import LinearRegression

# Build a Pipeline containing preprocessing and a regression model
model_pipeline = Pipeline(steps=[
    ('preprocessor', preprocessor),  # Data preprocessing steps
    ('regressor', LinearRegression())  # Regression model
])

# Train the model
model_pipeline.fit(X_train, y_train)

# Make predictions
y_pred = model_pipeline.predict(X_test)

# Output prediction results
print("\nPrediction results:")
print(y_pred)

The output result is:

预测结果:
[6375000.00000001 4874999.99999998]

4. Model Evaluation

In model evaluation, we usually use the mean squared error (MSE) and the coefficient of determination (R²) to evaluate the performance of a regression model.

Example

from sklearn.metrics import mean_squared_error, r2_score

# Calculate mean squared error (MSE) and coefficient of determination (R²)
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)

# Output evaluation results
print("\nModel evaluation:")
print(f"Mean Squared Error (MSE): {mse:.2f}")
print(f"Coefficient of Determination (R²): {r2:.2f}")

The output result is:

预测结果:
[6375000.00000001 4874999.99999998]
tianqixin@Mac-mini example-test % python3 test.py

模型评估:
均方误差 (MSE): 648125000000.03
决定系数 (R²): -63.81

5. Model Optimization (Grid Search)

To improve the model's performance, we can use GridSearchCV to tune the model's hyperparameters. In this example, we do not adjust the parameters of linear regression because it does not have many hyperparameters to tune. For more complex models (e.g., Random Forest, XGBoost, etc.), we can find the best parameters through grid search.

Example

from sklearn.model_selection import GridSearchCV

# 5. Model optimization: use grid search to tune hyperparameters
# Tune the hyperparameters of linear regression (only adjust 'fit_intercept')
param_grid = {
    'regressor__fit_intercept': [True, False],  # Whether to fit the intercept
}

grid_search = GridSearchCV(model_pipeline, param_grid, cv=5, scoring='neg_mean_squared_error', verbose=1)
grid_search.fit(X_train, y_train)

# Output the best parameters and results
print("\nBest parameters:")
print(grid_search.best_params_)

# Use the best model to make predictions
best_model = grid_search.best_estimator_
y_pred_optimized = best_model.predict(X_test)

# Output the optimized evaluation results
mse_opt = mean_squared_error(y_test, y_pred_optimized)
r2_opt = r2_score(y_test, y_pred_optimized)

print("\nOptimized model evaluation:")
print(f"Mean Squared Error (MSE): {mse_opt:.2f}")
print(f"Coefficient of Determination (R²): {r2_opt:.2f}")

The output result is:

Fitting 5 folds for each of 2 candidates, totalling 10 fits

最佳参数:
{'regressor__fit_intercept': True}

优化后的模型评估:
均方误差 (MSE): 648125000000.03
决定系数 (R²): -63.81

Complete Code

Complete code for Beijing house price prediction:

Example

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score

# Simulated data: contains house area, number of rooms, floor, year built, location (categorical variable), and house price (target variable)
data = {
    'area': [70, 85, 100, 120, 60, 150, 200, 80, 95, 110],
    'rooms': [2, 3, 3, 4, 2, 5, 6, 3, 3, 4],
    'floor': [5, 2, 8, 10, 3, 15, 18, 7, 9, 11],
    'year_built': [2005, 2010, 2012, 2015, 2000, 2018, 2020, 2008, 2011, 2016],
    'location': ['Chaoyang', 'Haidian', 'Chaoyang', 'Dongcheng', 'Fengtai', 'Haidian', 'Chaoyang', 'Fengtai', 'Dongcheng', 'Haidian'],
    'price': [5000000, 6000000, 6500000, 7000000, 4500000, 10000000, 12000000, 5500000, 6200000, 7500000]  # House price (target variable)
}

# Create DataFrame
df = pd.DataFrame(data)

# View data
print("Data preview:")
print(df.head())

# Feature selection
X = df[['area', 'rooms', 'floor', 'year_built', 'location']]  # Feature data
y = df['price']  # Target variable

# Split training set and test set
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Data preprocessing: standardize numerical features, One-Hot encode categorical features
numeric_features = ['area', 'rooms', 'floor', 'year_built']
categorical_features = ['location']

# Numerical feature preprocessing: standardization
numeric_transformer = Pipeline(steps=[
    ('scaler', StandardScaler())
])

# Categorical feature preprocessing: One-Hot encoding, set handle_unknown='ignore'
categorical_transformer = Pipeline(steps=[
    ('onehot', OneHotEncoder(handle_unknown='ignore'))  # Handle new categories in the test set
])

# Merge the processing steps for numerical and categorical features
preprocessor = ColumnTransformer(
    transformers=[
        ('num', numeric_transformer, numeric_features),
        ('cat', categorical_transformer, categorical_features)
    ]
)

# 3. Build the model
# Use a linear regression model, combined with data preprocessing steps
model_pipeline = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('regressor', LinearRegression())
])

# Train the model
model_pipeline.fit(X_train, y_train)

# Make predictions
y_pred = model_pipeline.predict(X_test)

# Output prediction results
print("\nPrediction results:")
print(y_pred)

# 4. Model evaluation: calculate mean squared error (MSE) and R² coefficient of determination
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)

print("\nModel evaluation:")
print(f"Mean Squared Error (MSE): {mse:.2f}")
print(f"Coefficient of Determination (R²): {r2:.2f}")

# 5. Model optimization: use grid search to tune hyperparameters
# Tune the hyperparameters of linear regression (only adjust 'fit_intercept')
param_grid = {
    'regressor__fit_intercept': [True, False],  # Whether to fit the intercept
}

grid_search = GridSearchCV(model_pipeline, param_grid, cv=5, scoring='neg_mean_squared_error', verbose=1)
grid_search.fit(X_train, y_train)

# Output the best parameters and results
print("\nBest parameters:")
print(grid_search.best_params_)

# Use the best model to make predictions
best_model = grid_search.best_estimator_
y_pred_optimized = best_model.predict(X_test)

# Output the optimized evaluation results
mse_opt = mean_squared_error(y_test, y_pred_optimized)
r2_opt = r2_score(y_test, y_pred_optimized)

print("\nOptimized model evaluation:")
print(f"Mean Squared Error (MSE): {mse_opt:.2f}")
print(f"Coefficient of Determination (R²): {r2_opt:.2f}")

After running this code, the output is:

数据预览:
   area  rooms  floor  year_built   location    price
0    70      2      5        2005   Chaoyang  5000000
1    85      3      2        2010    Haidian  6000000
2   100      3      8        2012   Chaoyang  6500000
3   120      4     10        2015  Dongcheng  7000000
4    60      2      3        2000    Fengtai  4500000

预测结果:
[6375000.00000001 4874999.99999998]

模型评估:
均方误差 (MSE): 648125000000.03
决定系数 (R²): -63.81
Fitting 5 folds for each of 2 candidates, totalling 10 fits

最佳参数:
{'regressor__fit_intercept': True}

优化后的模型评估:
均方误差 (MSE): 648125000000.03
决定系数 (R²): -63.81
Other Extensions