Gradient Descent -- Descending along the steepest direction
This chapter implements gradient descent from scratch and uses it to fit linear regression.This is the mathematical prototype of the entire AI training loop.
Concept Analysis
Core Formula
\[ \theta_{t+1} = \theta_t - \eta \nabla J(\theta_t) \]The Four-Step Training Loop
Forward Propagation
Predict Output
Compute Loss
Measure the Gap
Compute Gradient
Backpropagation
Update Parameters
θ -= η·∇J
Impact of Learning Rate
| Learning Rate | Effect |
|---|---|
| Too Small | Converges extremely slowly, requiring many steps |
| Moderate | Converges smoothly and quickly |
| Too Large | Oscillates or even diverges |
Real-life Example
Descending the Mountain on a Foggy Day
On a foggy day where you can't see your hand in front of your face, you need to descend the mountain. The strategy for each step:
Feel the steepest direction under your feet → Take a small step in the steepest downhill direction → Stop → Feel again → Take another small step → Repeat.
This is gradient descent: feeling = computing the gradient, taking a small step = updating parameters, stopping = next iteration.
Python Hands-On Practice
Example
# Generate data y = 3x + 2 + noise
np.random.seed(42)
X = np.linspace(0, 10, 100)
y = 3 * X + 2 + np.random.normal(0, 2, 100)
# Implement gradient descent from scratch
w, b = 0.0, 0.0
lr, n_iter = 0.01, 200
losses = []
for i in range(n_iter):
y_pred = w * X + b
loss = np.mean((y_pred - y) ** 2)
losses.append(loss)
# Gradient
dw = 2 * np.mean((y_pred - y) * X)
db = 2 * np.mean(y_pred - y)
# Update
w -= lr * dw
b -= lr * db
print(f"EXAMPLE gradient descent result:")
print(f"True: w=3.0, b=2.0")
print(f"Fitted: w={w:.4f}, b={b:.4f}")
print(fInitial loss: {losses[0]:.2f} → Final loss: {losses[-1]:.2f})
EXAMPLE 梯度下降结果: 真实: w=3.0, b=2.0 拟合: w=2.9874, b=2.2835 初始损失: 79.69 → 最终损失: 3.91
Loss Descent Curve
Application Scenarios in AI
The Mathematical Prototype of the PyTorch Training Loop
All deep learning training loops follow the same pattern, which is gradient descent:
① optimizer.zero_grad() → zero the gradient cache; ② loss = model(x) → forward propagation; ③ loss.backward() → backpropagation to compute gradients; ④ optimizer.step() → perform parameter update θ -= η·∇J.
Once you understand this four-step loop, you understand the core training logic of all deep learning frameworks.
Learning Rate Is the Most Important Hyperparameter
Learning rate too large → loss oscillates or even diverges; learning rate too small → convergence is extremely slow and may get stuck in local optima. Learning rate scheduling (Step, Cosine Annealing, Warmup) is an indispensable technique for training large models — GPT-3 training used a linear warmup + cosine decay strategy.
SGD Batch Size Trade-off
Full-batch GD: uses all data to compute exact gradients, slow but stable; SGD (batch_size=1): looks at only one sample at a time, fast but gradient estimates are noisy; Mini-batch SGD: a compromise (batch_size=32/64/128), with moderate noise, which actually helps escape local optima — this is the standard practice in real-world training.
Limitations of Gradient Descent
Gradient descent finds points where the gradient is zero — these may be global optima, local optima, or saddle points. In high-dimensional spaces, saddle points are far more common than local optima: a point is a minimum in some directions and a maximum in others. Adaptive optimizers such as Adam can escape saddle points more effectively.
Other extensions