PyTorch torch.nn.GELU Function

PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual


torch.nn.GELUIt is the Gaussian Error Linear Unit activation function in PyTorch.

It is the default activation function of the Transformer architecture, with better performance and smoother gradients compared to ReLU.

Function Definition

torch.nn.GELU(approximate='none')

Parameter Description:

  • approximate(str): Approximation algorithm. Optional'none'、'tanh'. Defaults to'none'。

Mathematical Principle

The mathematical formula of GELU:

GELU(x) = x * Φ(x)

where Φ(x) is the cumulative distribution function (CDF) of the standard normal distribution.

When using the tanh approximation:

GELU(x) ≈ 0.5x * (1 + tanh(√(2/π) * (x + 0.044715 * x³)))

Usage Examples

Example 1: Basic Usage

Create and use GELU activation:

Example

import torch
import torch.nn as nn

# Create GELU activation layer
gelu = nn.GELU()

# Test input
x = torch.tensor([-2.0, -1.0, 0.0, 1.0, 2.0])

# Forward pass
output = gelu(x)

print("Input:", x.tolist())
print("Output:", output.tolist())
print("nObservation: negative values have slight activation (non-zero), positive values continue to grow")

Example 2: Comparing Different Activation Functions

Compare GELU, ReLU, Sigmoid:

Example

import torch
import torch.nn as nn

x = torch.linspace(-4, 4, 21)

# Different activation functions
gelu = nn.GELU()
relu = nn.ReLU()
sigmoid = nn.Sigmoid()
tanh = nn.Tanh()

print("x       GELU      ReLU      Sigmoid   Tanh")
print("-" * 50)
for i in range(0, 21, 3):
    xi = x[i:i+3]
    print(f"{xi[0]:6.2f} {gelu(xi)[0]:8.4f} {relu(xi)[0]:8.4f} {sigmoid(xi)[0]:8.4f} {tanh(xi)[0]:8.4f}")

Example 3: Use in Transformer

Typical Transformer FFN layer:

h2 class="example">Example
import torch
import torch.nn as nn

class FeedForward(nn.Module):
    def __init__(self, d_model, dim_feedforward=2048, dropout=0.1):
        super(FeedForward, self).__init__()
        self.linear1 = nn.Linear(d_model, dim_feedforward)
        self.dropout = nn.Dropout(dropout)
        self.activation = nn.GELU()
        self.linear2 = nn.Linear(dim_feedforward, d_model)

    def forward(self, x):
        x = self.linear1(x)
        x = self.activation(x)
        x = self.dropout(x)
        x = self.linear2(x)
        return x

# Test FFN
ffn = FeedForward(d_model=512, dim_feedforward=2048)
x = torch.randn(32, 100, 512)  # (batch, seq, d_model)

output = ffn(x)
print("Input shape:", x.shape)
print("Output shape:", output.shape)

Example 4: Using tanh Approximation

Use tanh approximation to speed up computation:

Example

import torch
import torch.nn as nn

# Exact version
gelu_exact = nn.GELU(approximate='none')

# tanh approximation version
gelu_approx = nn.GELU(approximate='tanh')

x = torch.randn(1000)

output_exact = gelu_exact(x)
output_approx = gelu_approx(x)

# Compute difference
diff = (output_exact - output_approx).abs().max().item()
print(f"Max difference: {diff:.8f}")

# Performance comparison
import time

for _ in range(100):
    _ = gelu_exact(x)

start = time.time()
for _ in range(1000):
    _ = gelu_exact(x)
time_exact = time.time() - start

start = time.time()
for _ in range(1000):
    _ = gelu_approx(x)
time_approx = time.time() - start

print(f"Exact version time: {time_exact:.4f}s")
print(f"Approximate version time: {time_approx:.4f}s")

Activation Function Comparison

Activation Function Features Applicable Scenarios
nn.GELU Smooth, non-zero negative values, Transformer default Transformer、BERT、GPT
nn.ReLU Simple, sparse activation, dead neurons CNN, general deep learning
nn.SiLU Smooth, self-gating MobileNet、EfficientNet

Common Questions

Q1: What are the advantages of GELU compared to ReLU?

  • Negative values have slight activation, information is not lost
  • Smoother gradients, helpful for training
  • Better performance in Transformer

Q2: When to use the approximate version?

When inference speed is required and precision requirements are not strict, the tanh approximation is faster.

Q3: Can GELU be used in the output layer?

Usually not used in the output layer. Use Softmax for classification tasks, and the identity function for regression tasks.


Use Cases

nn.GELUMain application scenarios include:

  • Transformer architecture: Models such as BERT, GPT
  • Deep neural networks: Situations requiring smooth activation
  • Pretrained models: Modern NLP models

Tip: GELU is the most commonly used activation function in the NLP field and is standard equipment for Transformer.


PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual

Other Extensions