PyTorch torch.nn.ReLU Function
PyTorch torch.nn Reference Manual
torch.nn.ReLUIt is one of the most commonly used activation functions in PyTorch. It performs rectified linear unit (ReLU) operation element-wise on the input tensor.
ReLU is one of the most successful activation functions in deep learning because it is computationally simple, converges quickly, and effectively alleviates the vanishing gradient problem.
Function Definition
torch.nn.ReLU(inplace=False)
Parameter Description:
inplace(bool): If set toTrue, it will directly modify the input tensor in place, saving memory. Default isFalse。
Mathematical Principle
nn.ReLUThe calculation formula is as follows:
f(x) = max(0, x)
That is, for each element in the input, if the value is greater than 0, it remains unchanged; if the value is less than or equal to 0, the output is 0.
This simple nonlinear transformation allows the network to learn complex patterns while maintaining gradient flow.
Usage Examples
Example 1: Basic Usage
Create a ReLU activation layer and apply activation to the input:
Example
import torch.nn as nn
# Create ReLU activation layer
relu = nn.ReLU()
# Create an input tensor containing negative values
input_tensor = torch.tensor([[-1.0, 2.0, -3.0], [4.0, -5.0, 6.0]])
# Forward propagation
output = relu(input_tensor)
print("Input:\n", input_tensor)
print("Output:\n", output)
The output result is:
输入:
tensor([[-1., 2., -3.],
[ 4., -5., 6.]])
输出:
tensor([[0., 2., 0.],
[4., 0., 6.]])
As you can see, all negative values are set to 0, and positive values remain unchanged.
Example 2: In-place Operation
Using the inplace parameter can save memory:
Example
import torch.nn as nn
# Create a ReLU with in-place operation enabled
relu_inplace = nn.ReLU(inplace=True)
# Create input tensor
input_tensor = torch.randn(2, 4)
original_id = id(input_tensor)
# Modify in place
output = relu_inplace(input_tensor)
# Check whether the original tensor has been modified
print("Has the original tensor been modified:", id(output) == original_id)
print("Input/Output:\n", input_tensor)
The output result is:
是否修改了原张量: True
输入/输出:
tensor([[0.0000, 0.0000, 1.2345, 0.0000],
[0.5432, 0.0000, 0.0000, 2.3456]])
Note: In-place operations can save memory, but in some cases they may affect gradient computation. In the early stages of training or when intermediate activation values need to be preserved, it is recommended to use the default
inplace=False。
Example 3: Use in Neural Networks
In convolutional neural networks, ReLU usually follows the convolutional layer:
Example
import torch.nn as nn
# Define a simple convolutional neural network
class SimpleCNN(nn.Module):
def __init__(self):
super(SimpleCNN, self).__init__()
# Convolutional layer: input 3 channels, output 32 channels, kernel 3x3
self.conv1 = nn.Conv2d(3, 32, kernel_size=3, padding=1)
# ReLU activation
self.relu1 = nn.ReLU()
# Second convolutional layer
self.conv2 = nn.Conv2d(32, 64, kernel_size=3, padding=1)
self.relu2 = nn.ReLU()
# Pooling layer
self.pool = nn.MaxPool2d(2, 2)
def forward(self, x):
x = self.conv1(x)
x = self.relu1(x)
x = self.pool(x)
x = self.conv2(x)
x = self.relu2(x)
x = self.pool(x)
return x
# Create model
model = SimpleCNN()
# Print model structure
print("Model structure:")
print(model)
# Test forward propagation
# Simulate a 3x32x32 color image
input_image = torch.randn(1, 3, 32, 32)
output = model(input_image)
print("\nInput shape:", input_image.shape)
print("Output shape:", output.shape)
The output result is:
模型结构: SimpleCNN( (conv1): Conv2d(3, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (relu1): ReLU() (conv2): Conv2d(32, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (relu2): ReLU() (pool): MaxPool2d(kernel_size=2, stride=2, padding=0, dilation=1, ceil_mode=False) ) 输入形状: torch.Size([1, 3, 32, 32]) 输出形状: torch.Size([1, 64, 8, 8])
Example 4: Using nn.functional.relu
PyTorch also provides a functional interfacetorch.nn.functional.relu:
Example
import torch.nn.functional as F
# Using the functional interface
input_tensor = torch.randn(1, 5, 5)
# Method 1: Using the functional interface
output1 = F.relu(input_tensor)
# Method 2: Using the in-place version of the function
F.relu_(input_tensor)
print("Output:\n", output1)
The differences between the two:
nn.ReLUIs a module class that can be used in model definitions and saves parameters.F.reluIs a functional interface, commonly used in the forward method or in cases where there are no learnable parameters.F.relu_Is the underscore function of the in-place version.
Comparison with Other Activation Functions
| Activation Function | Formula | Characteristics | Applicable Scenarios |
|---|---|---|---|
nn.ReLU |
max(0, x) | Simple computation, sparse activation | General deep learning (default choice) |
nn.LeakyReLU |
x > 0 ? x : 0.01x | Small gradient on the negative axis | Prevents "dying neurons" | nn.GELU |
x * Φ(x) | Smooth approximation, default for Transformer | Transformer, BERT, etc. |
nn.Sigmoid |
1/(1+e^(-x)) | Output 0-1, gradient saturation | Output layer, binary classification |
nn.Tanh |
(e^x - e^(-x))/(e^x + e^(-x)) | Output -1 to 1, zero-centered | RNN、LSTM |
Frequently Asked Questions
Q1: Why does ReLU cause the "dying neuron" problem?
When the input remains negative, the output of ReLU is always 0, and the gradient is also 0, causing these neurons to be unable to learn further.
Solutions:
- Use
nn.LeakyReLUornn.ELUto replace - Use a smaller initial learning rate
- Adopt Batch Normalization
Q2: Is ReLU suitable for the output layer?
Usually not suitable. The output layer more commonly usesSigmoid(for classification) or the identity function (for regression), because ReLU's output is unbounded.
Use Cases
nn.ReLUIt is the most commonly used activation function in deep learning. The main application scenarios include:
- Convolutional Neural Networks: Introducing non-linearity after convolutional layers or fully connected layers.
- Multilayer Perceptron: As an activation function for hidden layers.
- Transformer: As an activation function in FFN (Feed-Forward Network) (although GELU is more commonly used now).
- Generative Adversarial Networks: As an activation function for the generator.
Tip: Unless there are special requirements, ReLU is the first choice for hidden layer activation functions when designing neural networks.
Other Extensions