Handwritten Gradient Descent to Fit a Straight Line
Without sklearn, use gradient descent to fit \(y = wx + b\) from scratch, observing step by step how the loss decreases.
After completing this case, you will understand:No closed-form solution is needed; as long as you can compute the gradient, the parameters will automatically converge to the optimal values.
Life Introduction
Basketball Shooting Practice — Adjust the Angle a Little Each Time
You stand outside the three-point line and shoot. The first shot goes left — next time adjust a bit to the right. The second shot goes right — adjust a bit to the left again. After each shot, you glance at the deviation and adjust your posture in the direction that reduces the deviation.
Machine learning training is exactly the same: each time after computing the predicted values, you look at the gap from the true values (the loss), then adjust the parameters in the direction that reduces the gap (gradient descent). Repeat a few hundred times, and the posture is set.
Intuitive Understanding
Our goal is to find a straight line \(y = wx + b\) that passes through all data points as closely as possible. \(w\) is the slope, \(b\) is the intercept, loss = the average of the squared vertical distances from all points to the line (MSE).
If we start from \(w=0, b=0\), the line is flat and the loss is large. Gradient descent will tell us: "w should be increased, and b should also be increased", and we follow. After repeating 100 times, the line will pass through the data points.
Mathematical Definition
\[ L(w, b) = \frac{1}{n}\sum_{i=1}^{n}(wx_i + b - y_i)^2 \] \[ \frac{\partial L}{\partial w} = \frac{2}{n}\sum_{i=1}^{n}(wx_i + b - y_i) \cdot x_i, \quad \frac{\partial L}{\partial b} = \frac{2}{n}\sum_{i=1}^{n}(wx_i + b - y_i) \] \[ w \leftarrow w - \eta \cdot \frac{\partial L}{\partial w}, \quad b \leftarrow b - \eta \cdot \frac{\partial L}{\partial b} \]Python Hands-on Practice
Example
np.random.seed(0)
# Generate noisy linear data: true y = 2x + 1
n = 60
x = np.random.uniform(-3, 3, n)
y = 2 * x + 1 + np.random.randn(60) * 0.8
# Manually derive the gradient function (take partial derivatives of w and b respectively)
def compute_gradients(w, b, x, y):
y_pred = w * x + b
error = y_pred - y # Predicted value - true value
dw = np.mean(2 * error * x) # dL/dw
db = np.mean(2 * error) # dL/db
return dw, db
# Gradient descent main loop
w, b = 0.0, 0.0 # Deliberately set initial parameters to 0
lr = 0.05
epochs = 100
print("EXAMPLE gradient descent process (true values w=2, b=1):")
for epoch in range(epochs):
y_pred = w * x + b
loss = np.mean((y_pred - y) ** 2)
dw, db = compute_gradients(w, b, x, y)
w -= lr * dw
b -= lr * db
if epoch % 20 == 0:
print(f" epoch {epoch:3d} loss={loss:.4f} w={w:.3f} b={b:.3f}")
print(f"\nEXAMPLE final: w={w:.3f}, b={b:.3f} (true w=2, b=1)")
print(f"Fitted line: y = {w:.2f}x + {b:.2f}")
EXAMPLE 梯度下降过程 (真实值 w=2, b=1): epoch 0 loss=13.4515 w=0.790 b=0.057 epoch 20 loss=0.9165 w=1.868 b=0.785 epoch 40 loss=0.7316 w=1.949 b=0.918 epoch 60 loss=0.7083 w=1.971 b=0.956 epoch 80 loss=0.7055 w=1.977 b=0.967 epoch 100 loss=0.7052 w=1.979 b=0.970 EXAMPLE 最终: w=1.979, b=0.970 (真实 w=2, b=1) 拟合直线: y = 1.98x + 0.97
Application Scenarios in AI
The code skeleton of this case is the template for all deep learning training: forward computation of predicted values → calculate loss → backpropagation to obtain gradients → update parameters → repeat. In PyTorch, optimizer.step() and loss.backward() automate this process.
Other extensions