Handwritten Gradient Descent to Fit a Straight Line

Without sklearn, use gradient descent to fit \(y = wx + b\) from scratch, observing step by step how the loss decreases.

After completing this case, you will understand:No closed-form solution is needed; as long as you can compute the gradient, the parameters will automatically converge to the optimal values.


Life Introduction

Basketball Shooting Practice — Adjust the Angle a Little Each Time

You stand outside the three-point line and shoot. The first shot goes left — next time adjust a bit to the right. The second shot goes right — adjust a bit to the left again. After each shot, you glance at the deviation and adjust your posture in the direction that reduces the deviation.

Machine learning training is exactly the same: each time after computing the predicted values, you look at the gap from the true values (the loss), then adjust the parameters in the direction that reduces the gap (gradient descent). Repeat a few hundred times, and the posture is set.


Intuitive Understanding

Our goal is to find a straight line \(y = wx + b\) that passes through all data points as closely as possible. \(w\) is the slope, \(b\) is the intercept, loss = the average of the squared vertical distances from all points to the line (MSE).

If we start from \(w=0, b=0\), the line is flat and the loss is large. Gradient descent will tell us: "w should be increased, and b should also be increased", and we follow. After repeating 100 times, the line will pass through the data points.


Mathematical Definition

\[ L(w, b) = \frac{1}{n}\sum_{i=1}^{n}(wx_i + b - y_i)^2 \] \[ \frac{\partial L}{\partial w} = \frac{2}{n}\sum_{i=1}^{n}(wx_i + b - y_i) \cdot x_i, \quad \frac{\partial L}{\partial b} = \frac{2}{n}\sum_{i=1}^{n}(wx_i + b - y_i) \] \[ w \leftarrow w - \eta \cdot \frac{\partial L}{\partial w}, \quad b \leftarrow b - \eta \cdot \frac{\partial L}{\partial b} \]

Python Hands-on Practice

Example

import numpy as np
np.random.seed(0)

# Generate noisy linear data: true y = 2x + 1
n = 60
x = np.random.uniform(-3, 3, n)
y = 2 * x + 1 + np.random.randn(60) * 0.8

# Manually derive the gradient function (take partial derivatives of w and b respectively)
def compute_gradients(w, b, x, y):
    y_pred = w * x + b
    error = y_pred - y          # Predicted value - true value
    dw = np.mean(2 * error * x) # dL/dw
    db = np.mean(2 * error)     # dL/db
    return dw, db

# Gradient descent main loop
w, b = 0.0, 0.0     # Deliberately set initial parameters to 0
lr = 0.05
epochs = 100

print("EXAMPLE gradient descent process (true values w=2, b=1):")
for epoch in range(epochs):
    y_pred = w * x + b
    loss = np.mean((y_pred - y) ** 2)
    dw, db = compute_gradients(w, b, x, y)
    w -= lr * dw
    b -= lr * db
    if epoch % 20 == 0:
        print(f"  epoch {epoch:3d}  loss={loss:.4f}  w={w:.3f}  b={b:.3f}")

print(f"\nEXAMPLE final: w={w:.3f}, b={b:.3f} (true w=2, b=1)")
print(f"Fitted line: y = {w:.2f}x + {b:.2f}")
EXAMPLE 梯度下降过程 (真实值 w=2, b=1):
  epoch   0  loss=13.4515  w=0.790  b=0.057
  epoch  20  loss=0.9165  w=1.868  b=0.785
  epoch  40  loss=0.7316  w=1.949  b=0.918
  epoch  60  loss=0.7083  w=1.971  b=0.956
  epoch  80  loss=0.7055  w=1.977  b=0.967
  epoch 100  loss=0.7052  w=1.979  b=0.970

EXAMPLE 最终: w=1.979, b=0.970  (真实 w=2, b=1)
拟合直线: y = 1.98x + 0.97

Application Scenarios in AI

The code skeleton of this case is the template for all deep learning training: forward computation of predicted values → calculate loss → backpropagation to obtain gradients → update parameters → repeat. In PyTorch, optimizer.step() and loss.backward() automate this process.

Other extensions