Machine Learning - House Price Prediction
When buying a house, people usually make a comprehensive judgment: location, area, property age, number of rooms, population density, transportation conditions, etc. Every factor influences the price, but this influence is not linear or singular; rather, it is the result of accumulation, trade-offs, and competition among factors.
In machine learning,regression problems, in essence, turn this "experience-based judgment" into a computable, reusable, and evaluable mathematical model.
This chapter will start from scratch and walk through the completehouse price predictionstandard machine learning pipeline: data understanding → feature analysis → model training → model evaluation → model optimization. The goal is not to memorize APIs, but to understand what each step does and why it is done.
Part 1: Project Preparation and Environment Setup
1.1 Core Tools Used
- NumPy: A low-level numerical computation tool that provides efficient array operations.
- Pandas: A core tool for tabular data analysis, essential in the early stages of machine learning.
- Matplotlib / Seaborn: Data visualization, used to understand data distributions and relationships.
- Scikit-learn: A machine learning toolbox covering datasets, models, and evaluation methods.
1.2 Installing Dependencies
pip install numpy pandas matplotlib seaborn scikit-learn
Part 2: Loading and Understanding the Dataset
Early tutorials often used the Boston Housing dataset, but it has been deprecated. Here we use the officially recommendedCalifornia Housingdataset, which has the same concept but more standardized data.
Example: Loading Data
from sklearn.datasets import fetch_california_housing
# Load the California housing price dataset
data = fetch_california_housing()
# Feature data (X)
df_features = pd.DataFrame(
data.data,
columns=data.feature_names
)
# Target variable (y): median house value
df_target = pd.DataFrame(
data.target,
columns=['MedHouseVal']
)
print(df_features.head())
print(df_target.head())
Output:
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude 0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88 -122.23 1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86 -122.22 2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85 -122.24 3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85 -122.25 4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85 -122.25 MedHouseVal 0 4.526 1 3.585 2 3.521 3 3.413 4 3.422
At this point, you should be clear on three things:
- Each row is the statistical features of one house
- Each column is a variable that can be used for prediction
MedHouseValis the target we want to predict
Part 3: Exploratory Data Analysis (EDA)
Before training the model, you must first answer one question:Is the data worth learning from?
3.1 Data Structure and Missing Value Check
Example: Data Overview
print("Target dimensions:", df_target.shape)
print("\n"Data types and missing values:")
df_features.info()
print("\n"Missing value statistics:")
print(df_features.isnull().sum())
Conclusion:
- Sufficient sample size (around 20,000 entries)
- All features are numerical
- No missing values, can directly build the model
3.2 Relationship Between a Single Feature and House Price
Machine learning is not a black box; at least at the introductory stage, you should know what the model is learning.
Example: Relationship Between Number of Rooms and House Price
plt.figure(figsize=(8, 6))
plt.scatter(df_features['AveRooms'], df_target['MedHouseVal'], alpha=0.4)
plt.xlabel('Average Rooms')
plt.ylabel('Median House Value')
plt.title('Rooms vs House Price')
plt.grid(True)
plt.show()
Intuitive conclusion: the more rooms, the higher the house price overall, but there is significant dispersion. This is exactly the purpose of a regression model.
3.3 Feature Correlation Analysis
Example: Correlation Heatmap
df_all = pd.concat([df_features, df_target], axis=1)
corr = df_all.corr()
plt.figure(figsize=(10, 8))
sns.heatmap(corr, cmap='coolwarm', center=0)
plt.title('Feature Correlation Heatmap')
plt.show()
The purpose of this step is not to "select a model," but to confirm: there is indeed a learnable statistical relationship.
Part 4: Building the First Regression Model
4.1 Splitting Training Set and Test Set
The model cannot use the same batch of data to both learn and be tested; otherwise, the evaluation results are meaningless.
Example: Dataset Splitting
X = df_features
y = df_target['MedHouseVal']
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
print("Training set:", X_train.shape)
print("Test set:", X_test.shape)
4.2 Training a Linear Regression Model
Example: Model Training
model = LinearRegression()
model.fit(X_train, y_train)
print("Intercept:", model.intercept_)
print("Coefficients:")
for name, coef in zip(X.columns, model.coef_):
print(f"{name}: {coef:.4f}")
The advantage of linear regression is:strong interpretability,suitable for understanding the essence of regression problems.
Part 5: Model Prediction and Evaluation
5.1 Comparison of Prediction Results
Example: Predictions vs. True Values
result = pd.DataFrame({
"Actual": y_test.values,
"Predicted": y_pred
})
print(result.head())
5.2 Quantifying the Model with Evaluation Metrics
Example: Evaluation Metrics
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print("MSE:", mse)
print("R2:", r2)
Explanation:
- The smaller the MSE, the lower the prediction error
- The closer R² is to 1, the stronger the model's explanatory power
Part 6: Model Optimization – Standardization and Ridge Regression
6.1 Why Standardize?
Huge differences in feature scales will cause the model to bias toward features with larger values.
Example: Feature Standardization
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
6.2 Using Ridge Regression to Suppress Overfitting
Example: Ridge Regression
ridge = Ridge(alpha=1.0)
ridge.fit(X_train_scaled, y_train)
y_pred_ridge = ridge.predict(X_test_scaled)
print("Ridge MSE:", mean_squared_error(y_test, y_pred_ridge))
print("Ridge R2:", r2_score(y_test, y_pred_ridge))
Part 7: Full Pipeline Review
