PyTorch Autograd Automatic Differentiation
The training of deep learning is essentially a process of repeatedly computing gradients and updating parameters.
Manually deriving the gradients for each layer is tedious and error-prone. PyTorch'sAutograd(automatic differentiation) engine was created to solve this problem—it canautomatically compute the gradients of any computation graph, allowing you to focus on model design rather than calculus derivation.
Core Concepts
1. What is Automatic Differentiation
Automatic Differentiation is not numerical differentiation (finite difference method), nor symbolic differentiation (algebraic derivation), but ratherrecording the computation process and applying the chain rule in reverse step by stepto compute derivatives precisely.
PyTorch's Autograd uses adynamic computation graph(Define-by-Run) approach: each forward pass constructs a directed acyclic graph (DAG) in real time, recording every operation and its inputs and outputs; during backpropagation, it traverses the graph in reverse and computes the gradient of each node in order.
2. requires_grad Attribute
The Tensor'srequires_gradattribute controls whether gradients need to be tracked for that tensor:
Example
# Create a tensor that requires gradient tracking (default requires_grad=False)
x = torch.tensor(3.0, requires_grad=True)
print(x) # tensor(3., requires_grad=True)
print(x.requires_grad) # True
# You can also modify it after creation
y = torch.tensor(2.0)
print(y.requires_grad) # False
y.requires_grad_(True) # In-place modification (note the trailing underscore)
print(y.requires_grad) # True
# Results of operations involving tensors with requires_grad=True automatically inherit requires_grad=True
z = x * y
print(z.requires_grad) # True
3. grad_fn and Computation Graph
Every tensor produced by an operation records agrad_fn, which points to the operation node that created it. This is the "skeleton" of the computation graph:
Example
x = torch.tensor(2.0, requires_grad=True)
y = torch.tensor(3.0, requires_grad=True)
z = x ** 2 + y * 3 # z = x² + 3y
print(z) # tensor(13., grad_fn=<AddBackward0>)
print(z.grad_fn) # <AddBackward0 object>
# Trace the chain of operations that created z
print(z.grad_fn.next_functions)
# ((<PowBackward0 object>, 0), (<MulBackward0 object>, 0))
# You can see that z is composed of a power operation and a multiplication operation
backward() Backpropagation
1. Calling backward() for Scalar Output
Call.backward()on the final scalar (loss value), and Autograd will automatically compute the gradients of all leaf nodes in reverse along the computation graph, storing the results in each tensor's.gradattribute:
Example
x = torch.tensor(2.0, requires_grad=True)
y = torch.tensor(3.0, requires_grad=True)
# Forward pass: z = x² + 3y
z = x ** 2 + y * 3
# Backward pass: automatically compute dz/dx and dz/dy
z.backward()
# View gradients
print(x.grad) # tensor(4.) ← dz/dx = 2x = 2×2 = 4
print(y.grad) # tensor(3.) ← dz/dy = 3
# Mathematical verification:
# z = x² + 3y
# dz/dx = 2x = 2×2 = 4 ✓
# dz/dy = 3 ✓
2. Accumulation Issue with Multiple backward() Calls
Autograd gradients areaccumulated, not overwritten. Each timebackward()is called, the gradients will be added to.gradthe existing values. This is the most common pitfall in training loops:
Example
x = torch.tensor(2.0, requires_grad=True)
# First backpropagation
loss = x ** 2
loss.backward()
print(x.grad) # tensor(4.) ← dL/dx = 2x = 4
# Second backpropagation (without zeroing!)
loss = x ** 2
loss.backward()
print(x.grad) # tensor(8.) ← Accumulated! Not 4, but 4+4=8
# ✅ Correct approach: zero the gradients before each backpropagation
x.grad.zero_() # Zero in-place (note the trailing underscore)
loss = x ** 2
loss.backward()
print(x.grad) # tensor(4.) ← Correct
When training a neural network, before each
backward()calloptimizer.zero_grad()to zero the gradients, otherwise gradients will keep accumulating and cause incorrect parameter updates.
3. Calling backward(gradient) for Non-Scalar Output
If the output is a vector or matrix rather than a scalar,backward()you need to pass agradientparameter with the same shape as the output (i.e., the "upstream gradient"), which is essentially computing the vector-Jacobian product (VJP):
Example
x = torch.tensor([1.0, 2.0, 3.0], requires_grad=True)
# Forward pass: y is a vector
y = x ** 2 # y = [1, 4, 9]
# For non-scalar output, you must pass the gradient parameter (shape same as y)
# The gradient can be understood as "the gradient of the loss with respect to y"
y.backward(gradient=torch.ones_like(y)) # Assume the upstream gradient is all 1s
print(x.grad)
# tensor([2., 4., 6.]) ← dy/dx = 2x, computed element-wise
# If the upstream gradient is not all 1s (e.g., weighted)
x.grad.zero_()
y.backward(gradient=torch.tensor([1.0, 0.5, 2.0])) # Different weights
# Actual computation: x.grad = 2x * gradient = [2×1, 4×0.5, 6×2]
print(x.grad)
# tensor([2., 2., 12.])
# More common approach: first use sum/mean to convert to a scalar, then call backward()
x.grad.zero_()
loss = (x ** 2).sum() # Aggregate the vector into a scalar
loss.backward()
print(x.grad)
# tensor([2., 4., 6.]) ← Equivalent to the first approach
torch.no_grad() Stopping Gradient Tracking
During model inference (prediction), there is no need to compute gradients. Usingtorch.no_grad()can skip the construction of the computation graph, significantly saving memory and computation:
Example
x = torch.tensor(3.0, requires_grad=True)
# In a no_grad context, no operations will track gradients
with torch.no_grad():
y = x ** 2
print(y.requires_grad) # False ← No longer tracking gradients
print(y.grad_fn) # None ← No computation graph node
# After exiting the no_grad context, normal tracking resumes
z = x ** 2
print(z.requires_grad) # True
# Common use: wrap the entire inference process during model evaluation
model = torch.nn.Linear(10, 1)
inputs = torch.randn(32, 10)
with torch.no_grad():
outputs = model(inputs) # No computation graph is built; faster speed and lower memory usage
@torch.no_grad() Decorator Syntax
You can also use the decorator form, which is suitable for marking an entire inference function as no-gradient:
Example
import torch.nn as nn
model = nn.Linear(10, 1)
@torch.no_grad()
def predict(model, x):
"""Inference function, no gradient computation needed"""
return model(x)
x = torch.randn(5, 10)
output = predict(model, x)
print(output.requires_grad) # False
detach() Separating from the Computation Graph
.detach()Returns a new tensor that shares data with the original tensor but does not track gradients. Commonly used in the following scenarios:
| Scenario | Description |
|---|---|
| Converting intermediate results to numpy arrays | numpy does not support tensors with gradients; you must detach() first |
| Recording training loss (logs) | Avoid keeping the entire computation graph to prevent memory leaks |
| Freezing gradient propagation for part of the network | Scenarios such as GAN training and transfer learning |
Example
x = torch.tensor([1.0, 2.0, 3.0], requires_grad=True)
y = x ** 2 + x * 3 # y has grad_fn
# The tensor after detach shares data with y but is detached from the computation graph
y_detached = y.detach()
print(y_detached.requires_grad) # False
print(y_detached.grad_fn) # None
# ✅ Convert to numpy (tensors with gradients cannot be directly converted)
# y.numpy() # ❌ Error: RuntimeError
y_detached.numpy() # ✅ Normal
# ✅ When recording loss values, detach (to avoid retaining the computation graph and consuming memory)
losses = []
for i in range(3):
loss = (x ** 2).sum()
losses.append(loss.detach().item()) # .item() converts a scalar tensor to a Python float
loss.backward()
x.grad.zero_()
print(losses) # [14.0, 14.0, 14.0]
retain_graph Retaining the Computation Graph
By default,backward()after execution, the computation graph will beautomatically released(to save memory). If you need to backpropagate multiple times on the same computation graph (e.g., in some GAN training), you need to passretain_graph=True:
Example
x = torch.tensor(2.0, requires_grad=True)
y = x ** 3 # y = x³
# First backward (retain computation graph)
y.backward(retain_graph=True)
print(x.grad) # tensor(12.) ← dy/dx = 3x² = 3×4 = 12
# Second backward (computation graph still exists)
x.grad.zero_()
y.backward(retain_graph=True)
print(x.grad) # tensor(12.) ← Same result
# No need to retain last time
x.grad.zero_()
y.backward() # After this, the computation graph is released
print(x.grad) # tensor(12.)
# Attempting backward again will raise an error (computation graph already released)
# y.backward() # ❌ RuntimeError: Trying to backward through the graph a second time
Unnecessary use
retain_graph=Truewill cause memory to keep growing, because the computation graph cannot be released. Only use it when multiple backward passes are actually needed.
Application of Gradients in Neural Network Training
The following is a complete example of manually implementing gradient descent using Autograd, demonstrating the full workflow of Autograd in actual training:
Example
# Construct training data: y = 2x + 1 plus noise
torch.manual_seed(42)
X = torch.randn(100, 1)
y_true = 2 * X + 1 + 0.1 * torch.randn(100, 1)
# Initialize model parameters (requires gradient tracking)
w = torch.zeros(1, requires_grad=True) # Weight
b = torch.zeros(1, requires_grad=True) # Bias
lr = 0.1 # Learning rate
epochs = 50 # Number of training epochs
for epoch in range(epochs):
# 1. Forward propagation: compute predictions
y_pred = X * w + b
# 2. Compute loss (mean squared error)
loss = ((y_pred - y_true) ** 2).mean()
# 3. Backward propagation: automatically compute d(loss)/dw and d(loss)/db
loss.backward()
# 4. Manually update parameters (wrap with no_grad to avoid update operations being tracked in the computation graph)
with torch.no_grad():
w -= lr * w.grad
b -= lr * b.grad
# 5. Clear gradients (must be cleared before the next backward)
w.grad.zero_()
b.grad.zero_()
if (epoch + 1) % 10 == 0:
print(f"Epoch {epoch+1:3d} | Loss: {loss.item():.4f} | w={w.item():.3f}, b={b.item():.3f}")
print(f"\nTraining complete: w ≈ {w.item():.3f} (true value 2.0), b ≈ {b.item():.3f} (true value 1.0))
The execution result of the above code is similar to the following:
Epoch 10 | Loss: 0.1064 | w=1.587, b=0.805 Epoch 20 | Loss: 0.0281 | w=1.876, b=0.949 Epoch 30 | Loss: 0.0152 | w=1.954, b=0.983 Epoch 40 | Loss: 0.0128 | w=1.978, b=0.993 Epoch 50 | Loss: 0.0122 | w=1.987, b=0.997 训练完成:w ≈ 1.987(真实值 2.0),b ≈ 0.997(真实值 1.0)
Common API Quick Reference Table
| API | Function | Common scenarios |
|---|---|---|
tensor.requires_grad_(True) |
Enable gradient tracking in-place | Enable Autograd for an existing tensor |
loss.backward() |
Backward propagation, compute gradients of all leaf nodes | Each iteration in the training loop |
tensor.grad |
Access the gradient value of a tensor | View or manually update parameters |
tensor.grad.zero_() |
Clear gradients in-place | Must clear before each backward() |
torch.no_grad() |
Context manager, disable gradient tracking | Inference phase, manual parameter updates |
tensor.detach() |
Return a new tensor detached from the computation graph | Convert to numpy, log records, freeze gradients |
tensor.item() |
Convert a scalar tensor to a Python number | Print loss values, record metrics |
loss.backward(retain_graph=True) |
Retain the computation graph, allow multiple backward passes | Scenarios such as GAN training that require multiple backward passes |
Common Questions and Precautions
1. Difference between leaf nodes and non-leaf nodes
Onlyleaf nodes(i.e., tensors directly created by the user, not the result of operations) have their gradients saved in.grad. Non-leaf nodes produced by intermediate operations do not retain gradients by default (to save memory). If you need to view the gradient of an intermediate node, call.retain_grad():
Example
x = torch.tensor(2.0, requires_grad=True) # Leaf node
# Intermediate node (non-leaf node)
y = x ** 2 # y is an intermediate node
y.retain_grad() # Explicitly declare to retain y's gradient
z = y * 3 # z is the final output
z.backward()
print(x.grad) # tensor(12.) ← Leaf node, saved normally
print(y.grad) # tensor(3.) ← Saved because of retain_grad()
# Intermediate nodes without retain_grad(), .grad is None
2. In-place operations may break the computation graph
Performing in-place operations on tensors that require gradient tracking (e.g.+=、.add_()) may prevent Autograd from correctly propagating backwards, so they should be avoided as much as possible:
Example
x = torch.tensor([1.0, 2.0], requires_grad=True)
# Dangerous: in-place operation on a leaf node
# x += 1 # May raise an error: a leaf Variable that requires grad has been used in an in-place operation
# Safe: use a non-in-place operation
y = x + 1 # Create a new tensor, do not modify x
# It is safe to perform in-place operations in a no_grad context (e.g., parameter updates)
with torch.no_grad():
x += 0.01 # Used for manual parameter updates, this is safe
3. Only floating-point tensors support gradients
Integer types (e.g.torch.int64) tensors do not supportrequires_grad=True, only floating-point types (float32、float64、float16) can participate in automatic differentiation.