Deep Learning Basics
Deep learning is not some mysterious black technology; it is essentially a mathematical function fitter—give it large amounts of input-output data, and it automatically finds the complex relationships between them.
For example: you give it a million pictures of cats, with each picture corresponding to the label "cat". Deep learning automatically figures out what combinations of pixels look like a cat.
Another example: you give it a billion sentences of human language, with the next word of each sentence as the target. Deep learning automatically learns, given the preceding context, what the next word is most likely to be.
The core advantage of deep learning is that as data volume increases, its performance continues to improve.This is something traditional machine learning cannot do.
In this chapter, we will start from scratch and break down every component of deep learning step by step:
- Mathematical Foundations of Neural Networks
- The Computational Process of Forward Propagation
- The Role of Activation Functions
- Loss Function Design
- The Principle of Backpropagation
- Optimizer Selection
- Regularization Techniques
- Normalization Methods
- Learning Rate Scheduling
Finally, we implement a complete multi-layer perceptron (MLP) with PyTorch, connecting all the knowledge points together.
This chapter involves some mathematics, but don't worry—we will explain every formula intuitively rather than asking you to memorize it. The focus is on understanding why it is designed this way, not how to calculate.
Mathematical Foundations of Neural Networks
Deep learning is built on three branches of mathematics: linear algebra, calculus, and probability theory.
You don't need to be a math expert, but you do need to understand a few core concepts.
Linear Algebra: Matrix Multiplication
The most basic operation in neural networks is matrix multiplication.
Why use matrices? Because they can concisely represent the process of "multiple inputs being transformed by multiple neurons."
Let's first look at an intuitive example:
Example
# Intuitive Understanding of Matrix Multiplication
# ============================================
import numpy as np
# Assume the input is a 3-dimensional vector: [height, weight, age]
# Units are: centimeters, kilograms, years
x = np.array([175, 70, 25]) # Input vector, shape (3,)
print(f"Input x: {x}")
print(f"Input shape: {x.shape}")
# Weight matrix W: 2 neurons, each neuron receives 3 inputs
# Shape is (output dimension, input dimension) = (2, 3)
W = np.array([
[0.1, 0.2, 0.3], # Weights of the 1st neuron
[0.4, 0.5, 0.6], # Weights of the 2nd neuron
])
print(f"\nWeight matrix W:\n{W}")
print(f"Weight shape: {W.shape}")
# Bias b: one bias per neuron
b = np.array([0.1, 0.2]) # Shape (2,)
print(f"\nBias b: {b}")
print(f"Bias shape: {b.shape}")
# Matrix multiplication: y = W · x + b
# Note: numpy's @ operator represents matrix multiplication
y = W @ x + b
print(f"\nOutput y: {y}")
print(f"Output shape: {y.shape}")
# Let's manually verify
y1 = 0.1 * 175 + 0.2 * 70 + 0.3 * 25 + 0.1 # The 1st neuron
y2 = 0.4 * 175 + 0.5 * 70 + 0.6 * 25 + 0.2 # The 2nd neuron
print(f"\nManual calculation verification: [{y1}, {y2}]")
The geometric meaning of matrix multiplication is a "linear transformation"—it can rotate, scale, and stretch vector spaces.
But linear transformations alone are not enough.
If every layer is just matrix multiplication, then no matter how deep the network is, the entire network is equivalent to a single-layer network, because the composition of multiple linear transformations is still a linear transformation.
This is why we need activation functions—to introduce nonlinearity.
Calculus: Derivatives and the Chain Rule
The core of training neural networks is "gradient descent"—finding the direction of parameters that minimizes the loss.
The gradient is the "derivative of a multivariable function"—it tells us how much the loss changes when each parameter changes a little.
Let's look at a simple example:
Example
# Intuitive understanding of derivatives
# ============================================
def f(x):
"""A simple function: f(x) = x²"""
return x ** 2
def numerical_derivative(f, x, h=1e-6):
"""Numerical derivative: use a small increment h to approximate the derivative
Definition of derivative: f'(x) = lim(h→0) [f(x+h) - f(x)] / h
"""
return (f(x + h) - f(x)) / h
# Compute the derivative at x=3
x = 3.0
df_dx = numerical_derivative(f, x)
print(f"f({x}) = {f(x)}")
print(f"f'({x}) ≈ {df_dx}")
print(f"Analytical solution (exact value): 2 * {x} = {2 * x}")
# Understand the meaning of the derivative:
# The derivative 6 means: at x=3, if x increases by 1, f(x) increases by about 6
x_new = x + 0.01
f_new = f(x_new)
print(f"\nx increases from {x} to {x_new}")
print(f"f(x) changes from {f(x)} to {f_new}")
print(f"Actual increase: {f_new - f(x)}")
print(f"Derivative prediction: {df_dx * 0.01}")
For multivariable functions, we need to compute the partial derivative with respect to each variable, then combine them into a vector—this is the gradient.
The chain rule is the mathematical foundation of backpropagation. It allows us to "break down layer by layer" the derivatives of composite functions.
Example
# Intuitive understanding of the chain rule
# ============================================
# Consider a composite function: y = f(g(x))
# Where: g(x) = x², f(z) = z³
# Then: y = (x²)³ = x⁶
def g(x):
return x ** 2
def f(z):
return z ** 3
def y(x):
return f(g(x))
# Manually compute the derivative (chain rule):
# dy/dx = df/dz * dz/dx = 3*z² * 2*x = 3*(x²)² * 2*x = 6*x⁵
x = 2.0
z = g(x) # z = 4
dy_dx_chain = 3 * (z ** 2) * 2 * x # Chain rule calculation
print(f"Chain rule calculation: dy/dx = {dy_dx_chain}")
# Numerical derivative verification
def numerical_derivative_y(x, h=1e-6):
return (y(x + h) - y(x)) / h
dy_dx_numerical = numerical_derivative_y(x)
print(f"Numerical derivative verification: dy/dx ≈ {dy_dx_numerical}")
# Analytical solution: 6*x^5 = 6*32 = 192
print(f"Analytical solution: 6 * {x}^5 = {6 * (x ** 5)}")
The core idea of the chain rule is:Break a complex function into simple functions, differentiate each, then multiply them together.。
Backpropagation is the application of the chain rule in neural networks—starting from the output, it propagates gradients backward layer by layer.
Probability Theory: Conditional Probability
In classification problems, neural networks often output a "probability distribution"—given an input, what is the probability of each class.
Conditional probability P(Y|X) represents "the probability of Y occurring given that X has occurred."
:.2%}
:.2%}
Example
# Use a neural network to output a probability distribution
# ============================================
import numpy as np
def softmax(x):
"""Softmax function: converts arbitrary real numbers into a probability distribution
The output values sum to 1, and each value is between [0, 1]
"""
# Subtract the maximum value to prevent numerical overflow
exp_x = np.exp(x - np.max(x))
return exp_x / np.sum(exp_x)
# Suppose the last layer of the neural network outputs 3 "scores" (logits)
# They correspond to "cat", "dog", "bird" respectively
logits = np.array([2.0, 1.0, 0.5])
print(fNeural network output (logits): {logits})
# Use Softmax to convert to probabilities
probabilities = softmax(logits)
print(fProbability distribution: {probabilities})
print(fSum of probabilities: {np.sum(probabilities)})
# Interpretation:
# P(cat|input image) ≈ 0.67
# P(dog|input image) ≈ 0.24
# P(bird|input image) ≈ 0.09
print("\nCategory probabilities:)
print(fCat: {probabilities)
print(fDog: {probabilities)
print(fBird: {probabilities)
Core points of three mathematical foundations:
| Mathematical branch | Core concept | Use in deep learning |
|---|---|---|
| Linear algebra | Matrix multiplication, vectors, tensors | Represent network structure, efficiently compute forward propagation |
| Calculus | Derivatives, chain rule, gradients | Backpropagation, update parameters |
| Probability theory | Conditional probability, probability distribution | Model uncertainty, design loss functions |
Forward Propagation
Forward propagation is the computation process where input data "flows through" the neural network, from the input layer to the hidden layer to the output layer.
Computation Graph
Representing a neural network as a computation graph makes the logic of forward propagation and backpropagation clearer.
A computation graph consists of nodes (operations) and edges (data flow).
Computation graph of a simple two-layer neural network:
Example
# Use a computation graph to understand forward propagation
# ============================================
import numpy as np
def relu(x):
"""ReLU activation function: max(0, x)"""
return np.maximum(0, x)
# A simple two-layer neural network
# Input layer (2) → hidden layer (3) → output layer (2)
# Input: x = [x1, x2]
x = np.array([1.0, 2.0])
print(fInput x: {x})
# First layer: input → hidden layer
W1 = np.array([
[0.1, 0.2], # 1st hidden neuron
[0.3, 0.4], # 2nd hidden neuron
[0.5, 0.6], # 3rd hidden neuron
])
b1 = np.array([0.1, 0.2, 0.3])
# Second layer: hidden layer → output layer
W2 = np.array([
[0.1, 0.2, 0.3], # 1st output neuron
[0.4, 0.5, 0.6], # 2nd output neuron
])
b2 = np.array([0.1, 0.2])
# Forward propagation computation graph
# Step 1: hidden layer linear transformation
z1 = W1 @ x + b1
print(f"\nHidden layer linear transformation z1: {z1})
# Step 2: hidden layer activation function
a1 = relu(z1)
print(fHidden layer activation a1: {a1})
# Step 3: output layer linear transformation
z2 = W2 @ a1 + b2
print(fOutput layer linear transformation z2: {z2})
# Step 4: output layer activation (for classification problems, usually Softmax)
a2 = softmax(z2) # Use the softmax function defined earlier
print(fOutput layer activation a2: {a2})
The advantage of the computation graph is that every step is clear, and during backpropagation you just compute gradients in the reverse order.
Neural Networks from the Perspective of Matrix Multiplication
From the perspective of matrix multiplication, a neural network is a series of "linear transformations + nonlinear activations".
Each layer can be abstracted as:
aᵢ = activation(Wᵢ · aᵢ₋₁ + bᵢ)
where a₀ is the input x.
This formula is concise yet powerful—it can represent almost all neural network structures, from simple logistic regression to complex Transformers.
Note the dimension matching for matrix multiplication: if the input vector is n-dimensional and the next layer has m neurons, then the weight matrix W must have shape m×n. This way, after W·x, you get an m-dimensional vector.
The Role of Activation Functions
As mentioned earlier: without activation functions, a deep network is equivalent to a single-layer network.
Use a simple example to prove this:
Example
# Why activation functions are needed
# ============================================
import numpy as np
# Suppose there is a two-layer network, but neither layer has an activation function
# Input
x = np.array([1.0, 2.0])
# First layer weights and bias
W1 = np.array([[0.1, 0.2], [0.3, 0.4]])
b1 = np.array([0.1, 0.2])
# Second layer weights and bias
W2 = np.array([[0.5, 0.6], [0.7, 0.8]])
b2 = np.array([0.3, 0.4])
# Compute the two layers separately
z1 = W1 @ x + b1
z2 = W2 @ z1 + b2
print(f"Two-layer calculation results: {z2}")
# Merge into a single layer for calculation
# It can be proved mathematically: z2 = (W2·W1)·x + (W2·b1 + b2)
W_combined = W2 @ W1
b_combined = W2 @ b1 + b2
z_combined = W_combined @ x + b_combined
print(f"Merged into one-layer calculation result: {z_combined}")
# The results are exactly the same!
print(f"\n"The two results are equal: {np.allclose(z2, z_combined)}")
print("Conclusion: without activation functions, deep network = single-layer network")
Activation functions are the key component that makes deep networks meaningful—they introduce nonlinearity and allow the network to learn complex patterns.
Activation Functions in Detail
An activation function determines how a neuron's output responds to its input.
A good activation function should satisfy several conditions:
- Nonlinearity: this is a must
- Differentiability: so gradients can be computed
- Simple computation: both forward and backward propagation should be fast
- No saturation: gradients will not vanish or explode
Let's look at the most commonly used activation functions.
Sigmoid: Saturation and Vanishing Gradients
Sigmoid is one of the earliest activation functions. It compresses any real number to between (0, 1).
Formula: σ(x) = 1 / (1 + e⁻ˣ)
Example
# Sigmoid activation function
# ============================================
import numpy as np
import matplotlib.pyplot as plt
def sigmoid(x):
"""Sigmoid function"""
return 1 / (1 + np.exp(-x))
def sigmoid_derivative(x):
"""Derivative of Sigmoid: σ'(x) = σ(x) * (1 - σ(x))"""
s = sigmoid(x)
return s * (1 - s)
# Test some values
x_values = [-10, -5, -2, 0, 2, 5, 10]
print("x | sigmoid(x) | sigmoid'(x)")
print("-" * 40)
for x in x_values:
s = sigmoid(x)
d = sigmoid_derivative(x)
print(f"{x:4} | {s:10.6f} | {d:12.6f}")
# Observation: when x is very large or very small, the derivative approaches 0
# This is the "vanishing gradient" problem
x_big = 10.0
print(f"\n"When x = {x_big}:")
print(f" sigmoid(x) = {sigmoid(x_big):.10f} (nearly 1)")
print(f" sigmoid'(x) = {sigmoid_derivative(x_big):.10f} (nearly 0)")
print("The gradient vanishes! During backpropagation, the gradient cannot be passed back.")
The problems with Sigmoid are obvious:
- When |x| > 6, the gradient is almost 0, causing vanishing gradients
- The output is not centered at 0, which affects optimization
- Computing the exponential is relatively slow
Due to these problems, Sigmoid is rarely used in hidden layers in modern deep learning.
ReLU and Variants
ReLU (Rectified Linear Unit) is the most commonly used activation function today.
Formula: ReLU(x) = max(0, x)
Simple but very effective.
Example
# ReLU and its variants
# ============================================
import numpy as np
def relu(x):
"""ReLU:max(0, x)"""
return np.maximum(0, x)
def relu_derivative(x):
"""Derivative of ReLU: 1 when x > 0, otherwise 0"""
return (x > 0).astype(float)
def leaky_relu(x, alpha=0.01):
"""Leaky ReLU: x when x > 0, otherwise alpha*x"""
return np.where(x > 0, x, alpha * x)
def gelu(x):
"""GELU:Gaussian Error Linear Unit
This is the activation function commonly used in Transformer
Approximation formula: 0.5 * x * (1 + tanh(sqrt(2/pi) * (x + 0.044715*x^3)))
"""
cdf = 0.5 * (1.0 + np.tanh(np.sqrt(2 / np.pi) * (x + 0.044715 * np.power(x, 3))))
return x * cdf
# Test
x_values = [-5, -2, -1, 0, 1, 2, 5]
print("x | ReLU | Leaky | GELU")
print("-" * 45)
for x in x_values:
r = relu(x)
lr = leaky_relu(x)
g = gelu(x)
print(f"{x:4} | {r:4.2f} | {lr:5.2f} | {g:5.2f}")
# Advantages of ReLU:
print("\n"Advantages of ReLU:")
print("1. Simple computation: only needs the max operation")
print("2. No saturation (positive interval): the gradient is always 1")
print("3. Sparsity: some neuron outputs are 0, increasing model robustness")
# Problem with ReLU: "dying ReLU"
print("\n"Problem with ReLU: dying ReLU")
print(If a neuron's input is always negative, then:)
print(- Output is always 0)
print(- Gradient is always 0)
print(- Parameters are never updated)
print(- This neuron is 'dead')
Comparison of several common activation functions:
| Activation function | Formula | Advantages | Disadvantages | Applicable scenarios |
|---|---|---|---|---|
| Sigmoid | 1/(1+e⁻ˣ) | Output in (0,1), interpretable as probability | Vanishing gradient, non-zero-centered, slow computation | Binary classification output layer (rarely used for hidden layers) |
| ReLU | max(0, x) | Fast computation, non-saturation, sparsity | Dead ReLU, non-zero-centered output | Most hidden layers (default choice) |
| Leaky ReLU | max(αx, x) | Solves the dead ReLU problem | α needs to be selected manually | Alternative when ReLU fails |
| GELU | x·Φ(x) | Smooth, works well in Transformer | Slightly more complex computation | Transformer, BERT, GPT, etc. |
| Swish/SwiGLU | x·σ(βx) | Superior to ReLU in deep networks | Complex computation | Large models such as PaLM, LLaMA |
Selection Principles
Recommended choices for activation functions:
- Hidden layer: Prefer ReLU, simple and effective
- If you encounter dead ReLU: Try Leaky ReLU or GELU
- Transformer: GELU or SwiGLU is the standard choice
- Binary classification output layer: Sigmoid (outputs probability)
- Multi-class classification output layer: Softmax (outputs probability distribution)
- Regression output layer: No activation (outputs any real number)
Don't overthink the choice of activation function. Usually ReLU is good enough. When you confirm ReLU is the bottleneck, it's not too late to switch to others.
Loss Functions
The loss function measures the gap between the "model's prediction" and the "true answer".
The goal of training is to make this gap as small as possible.
Cross-Entropy Loss (Classification)
The most commonly used loss function for classification problems is Cross-Entropy Loss.
Its intuitive meaning is: "how different is the probability distribution predicted by the model from the true distribution."
Example
# Cross-Entropy Loss
# ============================================
import numpy as np
def cross_entropy_loss(probs, target_one_hot):
"""Cross-Entropy Loss
probs: model-predicted probability distribution, shape (batch_size, num_classes)
target_one_hot: one-hot encoding of true labels, same shape as above
"""
# Add epsilon to prevent log(0)
epsilon = 1e-12
probs = np.clip(probs, epsilon, 1.0 - epsilon)
# cross_entropy = -sum(target * log(pred))
return -np.sum(target_one_hot * np.log(probs)) / len(probs)
def softmax(x):
"""Softmax function"""
exp_x = np.exp(x - np.max(x, axis=-1, keepdims=True))
return exp_x / np.sum(exp_x, axis=-1, keepdims=True)
# Example: three-class classification problem, classes are "cat", "dog", "bird"
# Assume the true label is "cat" (index 0), one-hot encoded
target = np.array([[1, 0, 0]]) # The correct answer is class 0
# Scenario 1: The model confidently predicts "cat"
logits1 = np.array([[5.0, 0.5, 0.1]]) # Cat has the highest score
probs1 = softmax(logits1)
loss1 = cross_entropy_loss(probs1, target)
print("Scenario 1: Model confidently predicts correctly")
print(f" Probability distribution: {probs1)
print(f" Loss: {loss1:.6f} (very small)")
# Scenario 2: The model is uncertain
logits2 = np.array([[1.0, 0.9, 0.8]]) # The scores of the three classes are similar
probs2 = softmax(logits2)
loss2 = cross_entropy_loss(probs2, target)
print("\n"Scenario 2: Model is uncertain")
print(f" Probability distribution: {probs2)
print(f" Loss: {loss2:.6f} (medium)")
# Scenario 3: The model predicts incorrectly
logits3 = np.array([[0.1, 5.0, 0.5]]) # Dog has the highest score
probs3 = softmax(logits3)
loss3 = cross_entropy_loss(probs3, target)
print("\n"Scenario 3: Model predicts incorrectly")
print(f" Probability distribution: {probs3)
print(f" Loss: {loss3:.6f} (very large)")
A simplified version of cross-entropy loss: when labels are class indices instead of one-hot encoding, a more efficient computation method can be used:
loss = -log(probs[target_class])
This is why in PyTorch, CrossEntropyLoss directly accepts class indices as labels.
MSE Loss (Regression)
For regression problems (predicting continuous values), the most commonly used loss is Mean Squared Error (MSE).
Formula: MSE = mean((pred - target)²)
Example
# MSE Loss
# ============================================
import numpy as np
def mse_loss(pred, target):
"""Mean Squared Error Loss"""
return np.mean((pred - target) ** 2)
def mae_loss(pred, target):
"""Mean Absolute Error Loss"""
return np.mean(np.abs(pred - target))
# Example: predicting house prices (unit: ten thousand yuan)
# Actual house prices
target = np.array([100, 200, 150])
# Scenario 1: Prediction is very accurate
pred1 = np.array([102, 195, 148])
mse1 = mse_loss(pred1, target)
mae1 = mae_loss(pred1, target)
print("Scenario 1: Accurate prediction")
print(f" Prediction: {pred1}")
print(f" Actual: {target}")
print(f" MSE: {mse1:.2f}")
print(f" MAE: {mae1:.2f}")
# Scenario 2: Prediction has a large error
pred2 = np.array([102, 280, 148]) # The second house prediction is 800,000 higher
mse2 = mse_loss(pred2, target)
mae2 = mae_loss(pred2, target)
print("\n"Scenario 2: There is a large error")
print(f" Prediction: {pred2}")
print(f" Actual: {target}")
print(f" MSE: {mse2:.2f} (greatly amplified by the large error)")
print(f" MAE: {mae2:.2f}")
# MSE vs MAE
print("\nMSE vs MAE:")
print("- MSE is more sensitive to large errors (because of squaring)")
print("- MAE is more robust to outliers")
print("- Generally, MSE is used more often")
Next-Word Prediction Loss for Language Models
The training objective of large language models (e.g., GPT) is simple: given the preceding context, predict the next word.
This is essentially a multi-class classification problem—each word in the vocabulary is a class.
Example
# Loss function for language models
# ============================================
import numpy as np
def language_model_loss(logits, targets, vocab_size):
"""Language model loss
Essentially, it is the cross-entropy loss at each position.
"""
batch_size, seq_len, _ = logits.shape
total_loss = 0.0
# Compute the loss at each position in the sequence
for i in range(seq_len):
# Predicted scores at this position
step_logits = logits[:, i, :]
# Convert to probabilities
step_probs = softmax(step_logits)
# Index of the target word
step_target = targets[:, i]
# Take only the probability of the target word
for b in range(batch_size):
target_idx = step_target[b]
target_prob = step_probs[b, target_idx]
# Loss = -log(p)
total_loss += -np.log(target_prob + 1e-12)
return total_loss / (batch_size * seq_len)
# A simple example
vocab_size = 10000 # Assume the vocabulary size is 10,000
batch_size = 2
seq_len = 3
# Model output: (batch_size, seq_len, vocab_size)
logits = np.random.randn(batch_size, seq_len, vocab_size)
# True targets: the word index that should be predicted at each position
targets = np.array([
[123, 456, 789], # Target word for the first sequence
[234, 567, 890], # Target word for the second sequence
])
loss = language_model_loss(logits, targets, vocab_size)
print(f"Language model loss: {loss:.4f}")
print("\n"The meaning of this loss:")
print("- For each position in the sequence")
print("- Predict the next word")
print("- Compute the average cross-entropy over all positions")
Summary of common loss functions:
| Task type | Loss function | Output layer activation | Label format |
|---|---|---|---|
| Binary classification | Binary Cross-Entropy | Sigmoid | 0 or 1 |
| Multi-class classification | Cross-Entropy | Softmax | Class index |
| Multi-label classification | Binary Cross-Entropy | Sigmoid | Multiple 0/1 |
| Regression | MSE / MAE | None | Continuous value |
| Language model | Cross-Entropy | Softmax | Word index |
Backpropagation
Backpropagation is the core algorithm for training neural networks.
Its goal is to compute the gradient of the loss function with respect to each parameter, and then use gradient descent to update the parameters.
Derivation via the Chain Rule
We will use a simple example to derive backpropagation step by step.
Example
# Manually implement backpropagation
# ============================================
import numpy as np
def relu(x):
return np.maximum(0, x)
def relu_derivative(x):
return (x > 0).astype(float)
def mse_loss(pred, target):
return np.mean((pred - target) ** 2)
# The simplest neural network: one linear transformation
# y = w * x + b
# Loss = (y - target)^2
# Input and target
x = np.array(2.0)
target = np.array(7.0)
# Initialize parameters
w = np.array(1.0)
b = np.array(0.0)
print(f"Initial parameters: w = {w}, b = {b}")
print(f"Target: {target}")
# ============================================
# Step 1: Forward propagation
# ============================================
y = w * x + b
loss = (y - target) ** 2
print(f"\n"Forward propagation:")
print(f" y = {y}")
print(f" loss = {loss}")
# ============================================
# Step 2: Backpropagation (manually compute gradients)
# ============================================
# We need to compute: d_loss/d_w, d_loss/d_b
# Use the chain rule step by step:
# loss = (y - target)^2
# d_loss/d_y = 2 * (y - target)
d_loss_d_y = 2 * (y - target)
# y = w * x + b
# d_y/d_w = x
# d_y/d_b = 1
d_y_d_w = x
d_y_d_b = 1.0
# Chain rule:
# d_loss/d_w = d_loss/d_y * d_y/d_w
# d_loss/d_b = d_loss/d_y * d_y/d_b
d_loss_d_w = d_loss_d_y * d_y_d_w
d_loss_d_b = d_loss_d_y * d_y_d_b
print(f"\n"Backpropagation:")
print(f" d_loss/d_y = {d_loss_d_y}")
print(f" d_loss/d_w = {d_loss_d_w}")
print(f" d_loss/d_b = {d_loss_d_b}")
# ============================================
# Step 3: Update parameters using gradient descent
# ============================================
learning_rate = 0.1
w_new = w - learning_rate * d_loss_d_w
b_new = b - learning_rate * d_loss_d_b
print(f"\n"Parameter update:")
print(f" w: {w} -> {w_new}")
print(f" b: {b} -> {b_new}")
# Verify: useNewParameterbeforetoward传播 # Verification: Forward propagation with new parameters
y_new = w_new * x + b_new
loss_new = (y_new - target) ** 2
print(f"\nVerification: Verification:)
print(f" 旧预测: {y}, 旧损失: {loss}" " Old prediction: {y}, old loss: {loss}")
print(f" New预测: {y_new}, New损失: {loss_new}" " New prediction: {y_new}, new loss: {loss_new}")
print(f" Loss decreased!" " Loss decreased!")
Although this example is simple, it contains all the core ideas of backpropagation:
- beforetoward传播:ComputeAllmiddleVariable Forward propagation: compute all intermediate variables
- 反toward传播: fromOutputstart,use链formula法ThenlayerlayerCompute梯degree Backpropagation: starting from the output, use the chain rule to compute gradients layer by layer
- ParameterUpdate:use梯degreeBottom降UpdateParameter Parameter update: update parameters using gradient descent
Gradient Flow in Computation Graphs
formulti-layer神经Networking,梯degreeWhen computingGraphMedium反towardflow动。 For multi-layer neural networks, gradients flow backward through the computational graph.
everyonelayerof梯degreeDependencyinafteronelayer传comeof梯degree。 The gradient of each layer depends on the gradient passed from the next layer.
Example
# twolayerNetworkingof反toward传播 # Backpropagation for a two-layer network
# ============================================
import numpy as np
def relu(x):
return np.maximum(0, x)
def relu_derivative(x):
return (x > 0).astype(float)
# A two-layer network # A two-layer network
# Input (2) → Hidden藏layer (3) → Output (1) # Input (2) → Hidden layer (3) → Output (1)
# Data # Data
x = np.array([1.0, 2.0]) # Input # Input
target = np.array([5.0]) # Target output # Target output
# Parameter initialization # Parameter initialization
W1 = np.array([[0.1, 0.2], [0.3, 0.4], [0.5, 0.6]]) # (3, 2)
b1 = np.array([0.1, 0.2, 0.3]) # (3,)
W2 = np.array([[0.7, 0.8, 0.9]]) # (1, 3)
b2 = np.array([0.4]) # (1,)
print(Initial parameters have been set.)
# ============================================
# Forward propagation # Forward propagation
# ============================================
print("\n=== Forward propagation ===" === Forward propagation ===)
# First layer # First layer
z1 = W1 @ x + b1 # (3,)
a1 = relu(z1) # (3,)
print(f"z1 = {z1}")
print(f"a1 = {a1}")
# Second layer # Second layer
z2 = W2 @ a1 + b2 # (1,)
a2 = z2 # RegressionTask, outputlayerNoaddactivate # Regression task, no activation on the output layer
print(f"z2 = {z2}")
print(f"a2 = {a2}")
# Loss # Loss
loss = np.mean((a2 - target) ** 2)
print(f"loss = {loss}")
# ============================================
# Backpropagation # Backpropagation
# ============================================
print("\n=== Backpropagation ===" === Backpropagation ===)
# Output layer gradient # Output layer gradient
d_loss_d_a2 = 2 * (a2 - target) # (1,)
d_a2_d_z2 = 1.0 # No activation function # No activation function
d_loss_d_z2 = d_loss_d_a2 * d_a2_d_z2 # (1,)
print(f"d_loss_d_z2 = {d_loss_d_z2}")
# No.二layerParameter梯degree # Second layer parameter gradients
d_z2_d_W2 = a1 # (3,), because z2 = W2@a1 + b2 # (3,), because z2 = W2@a1 + b2
d_z2_d_b2 = 1.0 # (1,)
d_z2_d_a1 = W2[0] # (3,)
# Compute parameter gradients # Compute parameter gradients
d_loss_d_W2 = d_loss_d_z2.reshape(-1, 1) @ d_z2_d_W2.reshape(1, -1) # (1, 3)
d_loss_d_b2 = d_loss_d_z2 * d_z2_d_b2 # (1,)
print(f"d_loss_d_W2 = {d_loss_d_W2}")
print(f"d_loss_d_b2 = {d_loss_d_b2}")
# Gradient passed to the hidden layer # Gradient passed to the hidden layer
d_loss_d_a1 = d_loss_d_z2 @ d_z2_d_a1 # (3,)
print(f"d_loss_d_a1 = {d_loss_d_a1}")
# Hidden layer gradient # Hidden layer gradient
d_a1_d_z1 = relu_derivative(z1) # (3,)
d_loss_d_z1 = d_loss_d_a1 * d_a1_d_z1 # (3,)
print(f"d_loss_d_z1 = {d_loss_d_z1}")
# First layer parameter gradients # First layer parameter gradients
d_z1_d_W1 = x # (2,)
d_z1_d_b1 = 1.0 # (3,)
# Compute parameter gradients # Compute parameter gradients
d_loss_d_W1 = d_loss_d_z1.reshape(-1, 1) @ d_z1_d_W1.reshape(1, -1) # (3, 2)
d_loss_d_b1 = d_loss_d_z1 * d_z1_d_b1 # (3,)
print(f"d_loss_d_W1 = {d_loss_d_W1}")
print(f"d_loss_d_b1 = {d_loss_d_b1}")
# ============================================
# Parameter update # Parameter update
# ============================================
print("\n=== Parameter update === === Parameter update ===)
learning_rate = 0.01
W1_new = W1 - learning_rate * d_loss_d_W1
b1_new = b1 - learning_rate * d_loss_d_b1
W2_new = W2 - learning_rate * d_loss_d_W2
b2_new = b2 - learning_rate * d_loss_d_b2
print("Parameters have been updated." "Parameters have been updated.")
ThisExamplesshows反toward传播ofCompleteProcess. This example demonstrates the complete process of backpropagation.
The core rule is: The core rules are:
- forLine性layer y = Wx + b:dL/dW = (dL/dy)·xᵀ,dL/db = dL/dy For a linear layer y = Wx + b: dL/dW = (dL/dy)·xᵀ, dL/db = dL/dy
- foractivateFunction y = f(z):dL/dz = dL/dy ⊙ f'(z)(⊙ Yes逐elementMultiplication) For an activation function y = f(z): dL/dz = dL/dy ⊙ f'(z) (⊙ is element-wise multiplication)
- 梯degreefromafter往before传,everyonelayeralluse链formula法Then"积累"梯degree Gradients propagate from back to front, and each layer uses the chain rule to "accumulate" gradients.
OKMessageYes:presentgenerationDeep LearningFramework(such as PyTorch、TensorFlow)willAutomaticisyouCompute反toward传播。you只requiresdefinitionbeforetoward传播,FrameworkKnow how to useAutomaticmicroDivide(autograd)搞decideAll梯degreeCalculate. The good news is: modern deep learning frameworks (such as PyTorch, TensorFlow) automatically compute backpropagation for you. You only need to define the forward pass, and the framework will handle all gradient computation with automatic differentiation (autograd).
Gradient Descent Optimizers
Compute出梯degreeAfter,I们requiresuseExcellenttransformdevicecomeUpdateparameters. After computing the gradients, we need an optimizer to update the parameters.
SimplestofExcellenttransformdeviceYesRandom梯degreeBottom降(SGD), but还Yes很many更AdvancedofExcellenttransformdevice. The simplest optimizer is stochastic gradient descent (SGD), but there are many more advanced optimizers.
SGD: Stochastic Gradient Descent
Formula: θ = θ - η·∇L(θ) Formula: θ = θ - η·∇L(θ)
where η is the learning rate. where η is the learning rate.
Example
# SGD optimizer # SGD optimizer
# ============================================
import numpy as np
def f(x):
"""I们want最smalltransformofFunction:f(x) = x²""" """The function we want to minimize: f(x) = x²"""
return x ** 2
def f_grad(x):
"""梯degree:f'(x) = 2x""" """Gradient: f'(x) = 2x"""
return 2 * x
# SGD optimization # SGD optimization
x = 10.0 # Initial value # Initial value
learning_rate = 0.1 # Learning rate # Learning rate
steps = 20
print(f"初start x = {x}, f(x) = {f(x)}" "Initial x = {x}, f(x) = {f(x)}")
print("-" * 40)
for step in range(steps):
grad = f_grad(x)
x = x - learning_rate * grad # SGD update # SGD update
print(f"Step {step+1}: x = {x:.6f}, f(x) = {f(x):.6f}")
print("-" * 40)
print(f"最终 x = {x:.6f}, f(x) = {f(x):.6f}" "Final x = {x:.6f}, f(x) = {f(x):.6f}")
print("manage论最Excellent:x = 0, f(x) = 0" "Theoretical optimum: x = 0, f(x) = 0")
SGD SimpleButYesDisadvantages: SGD is simple but has drawbacks:
- Uses the same learning rate for all parameters
- Convergence can be slow
- Easily gets stuck in local optima or saddle points.
- Sensitive to noise
Momentum: Acceleration with Momentum
Momentum 引入already"Speed"ofConcepts——Let梯degreeUpdate带Yes惯性。 Momentum introduces the concept of "velocity" — giving gradient updates inertia.
Formula: Formula:
v = β·v + (1-β)·∇L(θ) θ = θ - η·v
β is usually set to 0.9. β is usually set to 0.9.
Example
# Momentum optimizer # Momentum optimizer
# ============================================
import numpy as np
def f(x):
"""Objective function""" """Objective function"""
return x ** 2
def f_grad(x):
"""Gradient""" """Gradient"""
return 2 * x
def sgd_optimize(x_init, learning_rate, steps):
"""Pure SGD""" """Pure SGD"""
x = x_init
history = [x]
for _ in range(steps):
grad = f_grad(x)
x = x - learning_rate * grad
history.append(x)
return np.array(history)
def momentum_optimize(x_init, learning_rate, beta, steps):
"""Momentum"""
x = x_init
v = 0.0 # Initial velocity # Initial velocity
history = [x]
for _ in range(steps):
grad = f_grad(x)
v = beta * v + (1 - beta) * grad
x = x - learning_rate * v
history.append(x)
return np.array(history)
# Comparison # Comparison
x_init = 10.0
learning_rate = 0.1
beta = 0.9
steps = 20
sgd_history = sgd_optimize(x_init, learning_rate, steps)
momentum_history = momentum_optimize(x_init, learning_rate, beta, steps)
print("Step | SGD | Momentum")
print("-" * 35)
for i in range(steps + 1):
print(f"{i:4} | {sgd_history[i]:9.6f} | {momentum_history[i]:9.6f}")
print("\nObservation: Momentum converges faster!)
print(Because it accumulates past gradients, giving it inertia.)
Adam: Adaptive Learning Rate
Adam (Adaptive Moment Estimation) is currently one of the most commonly used optimizers.
It combines the ideas of Momentum (first moment) and RMSProp (second moment) to maintain an adaptive learning rate for each parameter.
Formula:
m = β₁·m + (1-β₁)·∇L(θ) v = β₂·v + (1-β₂)·(∇L(θ))² m̂ = m / (1-β₁ᵗ) v̂ = v / (1-β₂ᵗ) θ = θ - η·m̂ / (√v̂ + ε)
Default hyperparameters: β₁=0.9, β₂=0.999, ε=1e-8, η=1e-3.
Example
# Adam optimizer
# ============================================
import numpy as np
def adam_optimize(x_init, learning_rate, beta1, beta2, epsilon, steps):
"""Adam optimization"""
x = x_init
m = 0.0 # First moment
v = 0.0 # Second moment
history = [x]
for t in range(1, steps + 1):
grad = 2 * x # Gradient of f(x) = x²
# Update first and second moments
m = beta1 * m + (1 - beta1) * grad
v = beta2 * v + (1 - beta2) * (grad ** 2)
# Bias correction
m_hat = m / (1 - beta1 ** t)
v_hat = v / (1 - beta2 ** t)
# Parameter update
x = x - learning_rate * m_hat / (np.sqrt(v_hat) + epsilon)
history.append(x)
return np.array(history)
# Parameters
x_init = 10.0
learning_rate = 0.5
beta1 = 0.9
beta2 = 0.999
epsilon = 1e-8
steps = 20
# Optimization
history = adam_optimize(x_init, learning_rate, beta1, beta2, epsilon, steps)
print("Adam optimization process:")
print("Step | x | f(x)")
print("-" * 35)
for i in range(steps + 1):
print(f"{i:4} | {history[i]:9.6f} | {history[i]**2:9.6f}")
AdamW: Weight Decay Correction
AdamW is an improved version of Adam — it separates weight decay from the gradient.
This is useful for regularization.
Optimizer comparison:
| Optimizer | Formula | Advantages | Disadvantages | Applicable scenarios |
|---|---|---|---|---|
| SGD | θ = θ - η·∇L | Simple, stable | Slow convergence, sensitive to learning rate | Large datasets, rich experience in hyperparameter tuning |
| SGD + Momentum | v = βv + (1-β)∇L θ = θ - ηv | Accelerates convergence, reduces oscillation | Requires tuning β | Generally better than plain SGD |
| Adam | Adaptive learning rate | Fast convergence, requires less hyperparameter tuning | May generalize worse than SGD | Most scenarios (default choice) |
| AdamW | Adam + weight decay | Better regularization | Requires tuning weight decay coefficient | Transformers, large models |
Optimizer selection suggestion: Try Adam first, with the learning rate set to 1e-3 (3e-4 is a safer choice). If results are unsatisfactory, try tuning hyperparameters or switching to another optimizer.
Regularization Techniques
The goal of regularization is to prevent overfitting — helping the model perform well on unseen data.
L1/L2 Regularization
L1 and L2 regularization are the most basic regularization methods.
They add a penalty on parameters to the loss function:
L_total = L + λ·L_reg
L1 regularization (Lasso): L_reg = |w|
L2 regularization (Ridge): L_reg = w²
Example
# L1/L2 regularization
# ============================================
import numpy as np
def l1_regularization(weights, lambda_l1):
"""L1 regularization: returns the regularization term and gradient"""
reg = lambda_l1 * np.sum(np.abs(weights))
grad = lambda_l1 * np.sign(weights)
return reg, grad
def l2_regularization(weights, lambda_l2):
"""L2 regularization: returns the regularization term and gradient"""
reg = 0.5 * lambda_l2 * np.sum(weights ** 2)
grad = lambda_l2 * weights
return reg, grad
# Example
weights = np.array([1.0, -2.0, 3.0, 0.5])
print(f"Weights: {weights}")
lambda_l1 = 0.1
lambda_l2 = 0.1
l1_reg, l1_grad = l1_regularization(weights, lambda_l1)
print(f"\nL1 regularization term: {l1_reg:.4f}")
print(f"L1 gradient: {l1_grad}")
l2_reg, l2_grad = l2_regularization(weights, lambda_l2)
print(f"\nL2 regularization term: {l2_reg:.4f}")
print(f"L2 gradient: {l2_grad}")
print("\nL1 vs L2:")
print("- L1: tends to produce sparse solutions (many parameters become 0), can be used for feature selection")
print("- L2: tends to make all parameters very small, but not zero")
print("- In deep learning, L2 is more commonly used (or weight decay via AdamW)")
Dropout
Dropout is a simple but effective regularization technique: randomly "turning off" a portion of neurons during training.
This forces the model to learn more robust features — it cannot rely on any specific neuron.
Example
# Dropout
# ============================================
import numpy as np
def dropout_forward(x, dropout_rate, training=True):
"""Dropout forward pass"""
if not training:
# No dropout during testing, return directly
return x
# Generate mask: randomly set a portion of neurons to 0
mask = (np.random.rand(*x.shape) > dropout_rate).astype(float)
# Scaling: no scaling needed at test time, so divide by (1 - dropout_rate) during training
scale = 1.0 / (1.0 - dropout_rate)
# Apply dropout
out = x * mask * scale
return out, mask
# Test
x = np.array([1.0, 2.0, 3.0, 4.0, 5.0])
dropout_rate = 0.4
print(f"Input x: {x}")
print(f"Dropout rate: {dropout_rate}")
# Run multiple times to observe randomness
print("\nMultiple Dropout results:")
for i in range(3):
out, mask = dropout_forward(x, dropout_rate, training=True)
print(f" Round {i+1}: mask={mask}, out={out}")
# Test mode
out_test = dropout_forward(x, dropout_rate, training=False)
print(f"\nTest mode (no Dropout): {out_test}")
Data Augmentation
Data augmentation is one of the most effective regularization methods—it creates more training samples by transforming existing data.
For images, common data augmentation techniques include: flipping, rotation, scaling, cropping, color jittering, etc.
For text, common data augmentation techniques include: synonym replacement, random insertion, random deletion, back-translation, etc.
Summary of regularization techniques:
| Method | Principle | Advantages | Disadvantages |
|---|---|---|---|
| L2 Regularization | Penalizes large weights | Simple and stable | Limited effect |
| Dropout | Randomly drops neurons | Simple and effective | Slows down training |
| Data Augmentation | Increases diversity of training data | Significant effect | Requires domain knowledge |
| Early Stopping | Stop when validation set stops improving | Simple and effective | Requires monitoring the validation set |
| Batch Normalization | Normalize layer inputs | Accelerates training, regularizes | Adds computational overhead |
Batch Normalization vs Layer Normalization
Normalization techniques can speed up training and make optimization more stable.
Internal Covariate Shift
When training deep networks, the input distribution of each layer changes as the parameters of the previous layer change. This is "internal covariate shift" (Internal Covariate Shift).
This slows down training because each layer has to constantly adapt to new distributions.
The goal of normalization techniques is to keep the input distribution of each layer stable.
Batch Normalization
Batch Normalization (BatchNorm) normalizes along the batch dimension.
Formula:
μ = 1/N · sum(xᵢ) σ² = 1/N · sum((xᵢ - μ)²) x̂ᵢ = (xᵢ - μ) / √(σ² + ε) yᵢ = γ·x̂ᵢ + β
where γ and β are learnable parameters.
Example
# Batch Normalization
# ============================================
import numpy as np
def batch_norm(x, gamma, beta, epsilon=1e-5, training=True, running_mean=None, running_var=None, momentum=0.1):
"""Batch Normalization
x: input, shape (batch_size, features)
gamma: scaling parameter
beta: offset parameter
"""
if training:
# Training mode: use statistics of the current batch
batch_mean = np.mean(x, axis=0)
batch_var = np.var(x, axis=0)
# Update running statistics (for use at test time)
if running_mean is not None and running_var is not None:
running_mean[:] = momentum * running_mean + (1 - momentum) * batch_mean
running_var[:] = momentum * running_var + (1 - momentum) * batch_var
# Normalize
x_normalized = (x - batch_mean) / np.sqrt(batch_var + epsilon)
else:
# Test mode: use running statistics
x_normalized = (x - running_mean) / np.sqrt(running_var + epsilon)
# Scale and shift
out = gamma * x_normalized + beta
return out
# Test
batch_size = 4
features = 3
# One batch of data
x = np.array([
[1.0, 2.0, 3.0],
[2.0, 3.0, 4.0],
[3.0, 4.0, 5.0],
[4.0, 5.0, 6.0],
])
print(f"Input x:\n{x}")
# Initial parameters
gamma = np.ones(features) # Initially no scaling
beta = np.zeros(features) # Initially no offset
print(f"\ngamma: {gamma}")
print(f"beta: {beta}")
# Apply batch normalization
out = batch_norm(x, gamma, beta, training=True)
print(f"\nAfter batch normalization:\n{out}")
# Verify: mean of each feature should be close to 0, variance close to 1
print(f"\nNormalized mean: {np.mean(out, axis=0)}")
print(f"Normalized variance: {np.var(out, axis=0)}")
Advantages of batch normalization:
- Speeds up training
- Allows larger learning rates
- Reduces sensitivity to initialization
- Has a slight regularization effect
Disadvantages of batch normalization:
- Depends on batch size—works poorly with too-small batches
- Behavior differs between training and testing
- Not suitable for sequence models (e.g., RNN, Transformer)
Layer Normalization
Layer Normalization (LayerNorm) normalizes along the feature dimension instead of the batch dimension.
This makes it independent of batch size and very suitable for sequence models.
Example
# Layer Normalization
# ============================================
import numpy as np
def layer_norm(x, gamma, beta, epsilon=1e-5):
"""Layer Normalization
x: input, shape (batch_size, features) or (batch_size, seq_len, features)
"""
# Compute mean and variance over the last dimension (feature dimension)
mean = np.mean(x, axis=-1, keepdims=True)
var = np.var(x, axis=-1, keepdims=True)
# Normalize
x_normalized = (x - mean) / np.sqrt(var + epsilon)
# Scale and shift
out = gamma * x_normalized + beta
return out
# Test
batch_size = 2
seq_len = 3
features = 4
# Simulate input in Transformer: (batch_size, seq_len, features)
x = np.array([
[[1.0, 2.0, 3.0, 4.0],
[2.0, 3.0, 4.0, 5.0],
[3.0, 4.0, 5.0, 6.0]],
[[4.0, 5.0, 6.0, 7.0],
[5.0, 6.0, 7.0, 8.0],
[6.0, 7.0, 8.0, 9.0]],
])
print(f"Input x shape: {x.shape}")
# Initialize parameters
gamma = np.ones(features)
beta = np.zeros(features)
# Apply layer normalization
out = layer_norm(x, gamma, beta)
print(f"Shape after layer normalization: {out.shape}")
# Verify: mean at each position is close to 0, variance close to 1
mean_check = np.mean(out, axis=-1)
var_check = np.var(out, axis=-1)
print(f"\nMean at each position:\n{mean_check}")
print(f"Variance at each position:\n{var_check}")
Comparison of the two normalization methods:
| Method | Normalization dimension | Depends on batch | Applicable scenarios | Representative models |
|---|---|---|---|---|
| BatchNorm | batch dimension | Yes | CNN, fixed batch size | ResNet、VGG |
| LayerNorm | feature dimension | no | Transformer, RNN, variable-length sequences | GPT、BERT、LLaMA |
For Transformer and large language models, LayerNorm is the standard choice. It computes statistics independently at each position, does not depend on batch size, and does not need to maintain running statistics.
Learning Rate Scheduling
Learning rate is one of the most important hyperparameters — too large and it won't converge, too small and convergence is too slow.
Learning rate scheduling is dynamically adjusting the learning rate during training.
Warmup
In the early stage of training, use a small learning rate to "warm up" and let the model stabilize, then increase to the target learning rate.
This is especially important for Transformer.
Cosine Annealing
Cosine annealing gradually decreases the learning rate following the shape of the cosine function.
Example
# Learning rate scheduling
# ============================================
import numpy as np
def warmup_schedule(step, warmup_steps, base_lr):
"""Linear warmup schedule"""
if step < warmup_steps:
return base_lr * (step + 1) / warmup_steps
return base_lr
def cosine_annealing_schedule(step, total_steps, base_lr, min_lr=0.0):
"""Cosine annealing schedule"""
return min_lr + 0.5 * (base_lr - min_lr) * (1 + np.cos(np.pi * step / total_steps))
def warmup_cosine_schedule(step, warmup_steps, total_steps, base_lr, min_lr=0.0):
"""Warmup + cosine annealing"""
if step < warmup_steps:
return base_lr * (step + 1) / warmup_steps
return min_lr + 0.5 * (base_lr - min_lr) * (1 + np.cos(np.pi * (step - warmup_steps) / (total_steps - warmup_steps)))
# Parameters
total_steps = 100
warmup_steps = 10
base_lr = 1e-3
min_lr = 1e-5
print("Different learning rate schedules:")
print("Step | Warmup | Cosine | Warmup+Cosine")
print("-" * 55)
for step in range(0, total_steps + 1, 10):
lr_warmup = warmup_schedule(step, warmup_steps, base_lr)
lr_cosine = cosine_annealing_schedule(step, total_steps, base_lr, min_lr)
lr_combined = warmup_cosine_schedule(step, warmup_steps, total_steps, base_lr, min_lr)
print(f"{step:4} | {lr_warmup:.6f} | {lr_cosine:.6f} | {lr_combined:.6f}")
The Impact of Learning Rate on Training
The impact of learning rate:
- Too large: training unstable, may diverge
- Too small: slow convergence, may get stuck in local optima
- Appropriate: fast and stable convergence
A common strategy is the "learning rate finder": first train a few steps with learning rates increasing from small to large, and see at which learning rate the loss decreases fastest.
Summary of learning rate scheduling methods:
| Scheduling method | Characteristics | Applicable scenarios |
|---|---|---|
| Fixed learning rate | Simplest | Small models, simple tasks |
| Step Decay | Decrease every few steps | CNN, etc. |
| Cosine Annealing | Smooth decrease | Most scenarios |
| Warmup + Cosine | Increase first then decrease | Transformer, large models (recommended) |
| Reduce on Plateau | Decrease when validation set does not improve | Requires monitoring the validation set |
Hands-On: Implementing an MLP from Scratch with PyTorch
Now let's integrate all the previous knowledge points, use PyTorch to implement a complete multi-layer perceptron (MLP), and train it on a simple classification task.
Example
# Complete MLP training pipeline implemented in PyTorch
# ============================================
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import Dataset, DataLoader
import numpy as np
# ============================================
# 1. Prepare data
# ============================================
class SimpleDataset(Dataset):
"""A simple binary classification dataset"""
def __init__(self, n_samples=1000, n_features=10):
# Generate synthetic data
np.random.seed(42)
self.X = np.random.randn(n_samples, n_features).astype(np.float32)
# Simple classification rule: y = 1 if sum(x[:5]) > 0 else 0
self.y = (np.sum(self.X[:, :5], axis=1) > 0).astype(np.int64)
def __len__(self):
return len(self.X)
def __getitem__(self, idx):
return torch.tensor(self.X[idx]), torch.tensor(self.y[idx])
# Create dataset and data loader
train_dataset = SimpleDataset(n_samples=800)
val_dataset = SimpleDataset(n_samples=200)
train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True)
val_loader = DataLoader(val_dataset, batch_size=32, shuffle=False)
print(f"Training set size: {len(train_dataset)}")
print(f"Validation set size: {len(val_dataset)}")
# ============================================
# 2. Define the model
# ============================================
class MLP(nn.Module):
"""Multi-layer perceptron"""
def __init__(self, input_dim, hidden_dims, output_dim, dropout_rate=0.1):
super().__init__()
layers = []
prev_dim = input_dim
# Hidden layer
for hidden_dim in hidden_dims:
layers.append(nn.Linear(prev_dim, hidden_dim))
layers.append(nn.ReLU())
layers.append(nn.LayerNorm(hidden_dim)) # Layer normalization
layers.append(nn.Dropout(dropout_rate)) # Dropout
prev_dim = hidden_dim
# Output layer
layers.append(nn.Linear(prev_dim, output_dim))
self.model = nn.Sequential(*layers)
def forward(self, x):
return self.model(x)
# Create model
model = MLP(
input_dim=10,
hidden_dims=[64, 32],
output_dim=2,
dropout_rate=0.1
)
print("\nModel structure:")
print(model)
# ============================================
# 3. Define loss function and optimizer
# ============================================
criterion = nn.CrossEntropyLoss() # Cross-entropy loss
optimizer = optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-4) # AdamW optimizer
# Learning rate scheduler: warmup + cosine annealing
total_epochs = 20
warmup_epochs = 2
total_steps = len(train_loader) * total_epochs
warmup_steps = len(train_loader) * warmup_epochs
def lr_lambda(step):
if step < warmup_steps:
return (step + 1) / warmup_steps
else:
progress = (step - warmup_steps) / (total_steps - warmup_steps)
return 0.5 * (1 + np.cos(np.pi * progress))
scheduler = optim.lr_scheduler.LambdaLR(optimizer, lr_lambda=lr_lambda)
print("\nTraining configuration:")
print(f" Epochs: {total_epochs}")
print(f" Warmup epochs: {warmup_epochs}")
print(f" Total steps: {total_steps}")
print(f" Warmup steps: {warmup_steps}")
# ============================================
# 4. Training loop
# ============================================
print("\nStart training...")
print("-" * 50)
global_step = 0
best_val_acc = 0.0
for epoch in range(total_epochs):
# Training phase
model.train()
train_loss = 0.0
train_correct = 0
train_total = 0
for batch_X, batch_y in train_loader:
# Forward pass
logits = model(batch_X)
loss = criterion(logits, batch_y)
# Backward pass
optimizer.zero_grad()
loss.backward()
optimizer.step()
scheduler.step()
# Statistics
train_loss += loss.item()
predicted = torch.argmax(logits, dim=1)
train_total += batch_y.size(0)
train_correct += (predicted == batch_y).sum().item()
global_step += 1
train_loss /= len(train_loader)
train_acc = train_correct / train_total
# Validation phase
model.eval()
val_loss = 0.0
val_correct = 0
val_total = 0
with torch.no_grad():
for batch_X, batch_y in val_loader:
logits = model(batch_X)
loss = criterion(logits, batch_y)
val_loss += loss.item()
predicted = torch.argmax(logits, dim=1)
val_total += batch_y.size(0)
val_correct += (predicted == batch_y).sum().item()
val_loss /= len(val_loader)
val_acc = val_correct / val_total
# Save the best model
if val_acc > best_val_acc:
best_val_acc = val_acc
# You can save model checkpoint here
# Print log
current_lr = optimizer.param_groups[0]['lr']
print(f"Epoch {epoch+1:2d} | lr: {current_lr:.6f} | "
f"Train Loss: {train_loss:.4f} | Train Acc: {train_acc:.4f} | "
f"Val Loss: {val_loss:.4f} | Val Acc: {val_acc:.4f}")
print("-" * 50)
print(f"Training complete! Best validation accuracy: {best_val_acc:.4f}")
# ============================================
# 5. Inference example
# ============================================
print("\nInference example:")
model.eval()
# A few test samples
test_samples = [
np.array([1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0], dtype=np.float32), # Should be 1
np.array([-1.0, -1.0, -1.0, -1.0, -1.0, 0.0, 0.0, 0.0, 0.0, 0.0], dtype=np.float32), # Should be 0
]
for i, sample in enumerate(test_samples):
with torch.no_grad():
input_tensor = torch.tensor(sample).unsqueeze(0)
logits = model(input_tensor)
probs = torch.softmax(logits, dim=1)
predicted_class = torch.argmax(probs, dim=1).item()
print(f" Sample {i+1}: {sample[:5]}...")
print(f" Predicted class: {predicted_class}")
print(f" Probabilities: class 0={probs[0,0]:.4f}, class 1={probs[0,1]:.4f}")
This example contains all standard components of deep learning training:
- Data processing: Dataset, DataLoader
- Model definition: MLP with LayerNorm and Dropout
- Loss function: CrossEntropyLoss
- Optimizer: AdamW
- Learning rate scheduler: Warmup + Cosine
- Training loop: training phase + validation phase
- Model saving and inference
This workflow can be directly extended to more complex tasks and models.
Other extensions