PyTorch torch.nn.ReLU Function

PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual


torch.nn.ReLUIt is one of the most commonly used activation functions in PyTorch. It performs rectified linear unit (ReLU) operation element-wise on the input tensor.

ReLU is one of the most successful activation functions in deep learning because it is computationally simple, converges quickly, and effectively alleviates the vanishing gradient problem.

Function Definition

torch.nn.ReLU(inplace=False)

Parameter Description:

  • inplace(bool): If set toTrue, it will directly modify the input tensor in place, saving memory. Default isFalse。

Mathematical Principle

nn.ReLUThe calculation formula is as follows:

f(x) = max(0, x)

That is, for each element in the input, if the value is greater than 0, it remains unchanged; if the value is less than or equal to 0, the output is 0.

This simple nonlinear transformation allows the network to learn complex patterns while maintaining gradient flow.


Usage Examples

Example 1: Basic Usage

Create a ReLU activation layer and apply activation to the input:

Example

import torch
import torch.nn as nn

# Create ReLU activation layer
relu = nn.ReLU()

# Create an input tensor containing negative values
input_tensor = torch.tensor([[-1.0, 2.0, -3.0], [4.0, -5.0, 6.0]])

# Forward propagation
output = relu(input_tensor)

print("Input:\n", input_tensor)
print("Output:\n", output)

The output result is:

输入:
 tensor([[-1.,  2., -3.],
        [ 4., -5.,  6.]])
输出:
 tensor([[0., 2., 0.],
        [4., 0., 6.]])

As you can see, all negative values are set to 0, and positive values remain unchanged.

Example 2: In-place Operation

Using the inplace parameter can save memory:

Example

import torch
import torch.nn as nn

# Create a ReLU with in-place operation enabled
relu_inplace = nn.ReLU(inplace=True)

# Create input tensor
input_tensor = torch.randn(2, 4)
original_id = id(input_tensor)

# Modify in place
output = relu_inplace(input_tensor)

# Check whether the original tensor has been modified
print("Has the original tensor been modified:", id(output) == original_id)
print("Input/Output:\n", input_tensor)

The output result is:

是否修改了原张量: True
输入/输出:
 tensor([[0.0000, 0.0000, 1.2345, 0.0000],
        [0.5432, 0.0000, 0.0000, 2.3456]])

Note: In-place operations can save memory, but in some cases they may affect gradient computation. In the early stages of training or when intermediate activation values need to be preserved, it is recommended to use the defaultinplace=False。

Example 3: Use in Neural Networks

In convolutional neural networks, ReLU usually follows the convolutional layer:

Example

import torch
import torch.nn as nn

# Define a simple convolutional neural network
class SimpleCNN(nn.Module):
    def __init__(self):
        super(SimpleCNN, self).__init__()
        # Convolutional layer: input 3 channels, output 32 channels, kernel 3x3
        self.conv1 = nn.Conv2d(3, 32, kernel_size=3, padding=1)
        # ReLU activation
        self.relu1 = nn.ReLU()
        # Second convolutional layer
        self.conv2 = nn.Conv2d(32, 64, kernel_size=3, padding=1)
        self.relu2 = nn.ReLU()
        # Pooling layer
        self.pool = nn.MaxPool2d(2, 2)

    def forward(self, x):
        x = self.conv1(x)
        x = self.relu1(x)
        x = self.pool(x)

        x = self.conv2(x)
        x = self.relu2(x)
        x = self.pool(x)

        return x

# Create model
model = SimpleCNN()

# Print model structure
print("Model structure:")
print(model)

# Test forward propagation
# Simulate a 3x32x32 color image
input_image = torch.randn(1, 3, 32, 32)
output = model(input_image)

print("\nInput shape:", input_image.shape)
print("Output shape:", output.shape)

The output result is:

模型结构:
SimpleCNN(
  (conv1): Conv2d(3, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
  (relu1): ReLU()
  (conv2): Conv2d(32, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
  (relu2): ReLU()
  (pool): MaxPool2d(kernel_size=2, stride=2, padding=0, dilation=1, ceil_mode=False)
)

输入形状: torch.Size([1, 3, 32, 32])
输出形状: torch.Size([1, 64, 8, 8])

Example 4: Using nn.functional.relu

PyTorch also provides a functional interfacetorch.nn.functional.relu:

Example

import torch
import torch.nn.functional as F

# Using the functional interface
input_tensor = torch.randn(1, 5, 5)

# Method 1: Using the functional interface
output1 = F.relu(input_tensor)

# Method 2: Using the in-place version of the function
F.relu_(input_tensor)

print("Output:\n", output1)

The differences between the two:

  • nn.ReLUIs a module class that can be used in model definitions and saves parameters.
  • F.reluIs a functional interface, commonly used in the forward method or in cases where there are no learnable parameters.
  • F.relu_Is the underscore function of the in-place version.

Comparison with Other Activation Functions

Activation Function Formula Characteristics Applicable Scenarios
nn.ReLU max(0, x) Simple computation, sparse activation General deep learning (default choice)
nn.LeakyReLU x > 0 ? x : 0.01x Small gradient on the negative axis Prevents "dying neurons"
nn.GELU x * Φ(x) Smooth approximation, default for Transformer Transformer, BERT, etc.
nn.Sigmoid 1/(1+e^(-x)) Output 0-1, gradient saturation Output layer, binary classification
nn.Tanh (e^x - e^(-x))/(e^x + e^(-x)) Output -1 to 1, zero-centered RNN、LSTM

Frequently Asked Questions

Q1: Why does ReLU cause the "dying neuron" problem?

When the input remains negative, the output of ReLU is always 0, and the gradient is also 0, causing these neurons to be unable to learn further.

Solutions:

  • Usenn.LeakyReLUornn.ELUto replace
  • Use a smaller initial learning rate
  • Adopt Batch Normalization

Q2: Is ReLU suitable for the output layer?

Usually not suitable. The output layer more commonly usesSigmoid(for classification) or the identity function (for regression), because ReLU's output is unbounded.


Use Cases

nn.ReLUIt is the most commonly used activation function in deep learning. The main application scenarios include:

  • Convolutional Neural Networks: Introducing non-linearity after convolutional layers or fully connected layers.
  • Multilayer Perceptron: As an activation function for hidden layers.
  • Transformer: As an activation function in FFN (Feed-Forward Network) (although GELU is more commonly used now).
  • Generative Adversarial Networks: As an activation function for the generator.

Tip: Unless there are special requirements, ReLU is the first choice for hidden layer activation functions when designing neural networks.


PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual

Other Extensions