PyTorch torch.nn.Conv2d Function

PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual


torch.nn.Conv2dIt is a module in PyTorch for two-dimensional convolution, and is a core component of convolutional neural networks (CNN).

It extracts spatial features by applying learnable convolution kernels to the input tensor, and is widely used in image processing and computer vision tasks.

Function Definition

torch.nn.Conv2d(in_channels, out_channels, kernel_size, stride=1, padding=0, dilation=1, groups=1, bias=True, padding_mode='zeros')

Parameters:

  • in_channels(int): Number of input channels. For example, 3 for an RGB image.
  • out_channels(int): Number of output channels, i.e., the number of convolution kernels.
  • kernel_size(int or tuple): Size of the convolution kernel. Can be an integer (square) or a tuple (height x width).
  • stride(int or tuple): The stride of the convolution kernel movement. Default is 1.
  • padding(int or tuple): Padding size at the edges of the input. Default is 0.
  • dilation(int or tuple): Spacing between elements of the convolution kernel. Default is 1 (standard convolution).
  • groups(int): Number of groups for grouped convolution. Default is 1 (standard convolution).
  • bias(bool): Whether to add a bias term. Default isTrue。
  • padding_mode(str): Padding mode. Optional'zeros'、'reflect'、'replicate'、'circular'。

Attributes:

  • weight(Tensor): Learnable weights with shape (out_channels, in_channels/groups, kernel_size
  • bias(Tensor): Learnable bias with shape (out_channels,).

Usage Examples

Example 1: Basic Usage

Create a simple 2D convolution layer:

Example

import torch
import torch.nn as nn

# Create a convolution layer: input 3 channels, output 32 channels, kernel 3x3
conv = nn.Conv2d(in_channels=3, out_channels=32, kernel_size=3)

# Print the shapes of weights and bias
print("Weight shape:", conv.weight.shape)  # torch.Size([32, 3, 3, 3])
print("Bias shape:", conv.bias.shape)    # torch.Size([32])

# Create an input tensor: batch=1, channels=3, height=32, width=32
input_tensor = torch.randn(1, 3, 32, 32)

# Forward propagation
output = conv(input_tensor)

print("Input shape:", input_tensor.shape)  # torch.Size([1, 3, 32, 32])
print("Output shape:", output.shape)      # torch.Size([1, 32, 30, 30])

The output result is:

权重形状: torch.Size([32, 3, 3, 3])
偏置形状: torch.Size([32])
输入形状: torch.Size([1, 3, 32, 32])
输出形状: torch.Size([1, 32, 30, 30])

By default, padding=0, so the output size will decrease. If you need to keep the size, you can add padding.

Example 2: Using Padding to Keep Dimensions

By adding padding, the input and output sizes can be kept consistent:

Example

import torch
import torch.nn as nn

# Create a convolution layer with padding: padding=1 keeps the size
conv_pad = nn.Conv2d(in_channels=3, out_channels=32, kernel_size=3, padding=1)

# Input
input_tensor = torch.randn(1, 3, 32, 32)

# Forward propagation
output = conv_pad(input_tensor)

print("Input shape:", input_tensor.shape)
print("Output shape:", output.shape)  # Keep 32x32

The output result is:

输入形状: torch.Size([1, 3, 32, 32])
输出形状: torch.Size([1, 32, 32, 32])

Example 3: Different Stride and Dilation

Adjusting stride and dilation can change the output size and receptive field:

Example

import torch
import torch.nn as nn

# Strided convolution: stride=2 reduces the size
conv_stride = nn.Conv2d(3, 32, kernel_size=3, stride=2, padding=1)
input_tensor = torch.randn(1, 3, 32, 32)
output_stride = conv_stride(input_tensor)
print("Stride=2 -> Output shape:", output_stride.shape)

# Dilated convolution: dilation=2 increases the receptive field
conv_dilation = nn.Conv2d(3, 32, kernel_size=3, dilation=2)
output_dilation = conv_dilation(input_tensor)
print("Dilation=2 -> Output shape:", output_dilation.shape)

The output result is:

Stride=2 -> 输出形状: torch.Size([1, 32, 16, 16])
Dilation=2 -> 输出形状: torch.Size([1, 32, 28, 28])

Example 4: Grouped Convolution

The groups parameter can implement grouped convolution, commonly used in lightweight networks:

Example

import torch
import torch.nn as nn

# Grouped convolution: groups=2 divides the input into 2 groups
conv_group = nn.Conv2d(in_channels=4, out_channels=8, kernel_size=3, groups=2)

# Input with 4 channels
input_tensor = torch.randn(1, 4, 16, 16)

# Forward propagation
output = conv_group(input_tensor)

print("Input shape:", input_tensor.shape)
print("Output shape:", output.shape)
print("Weight shape:", conv_group.weight.shape)  # The weight shape is different after grouping

The output result is:

输入形状: torch.Size([1, 4, 16, 16])
输出形状: torch.Size([1, 8, 14, 14])
权重形状: torch.Size([8, 2, 3, 3])

Example 5: Use in Neural Networks

Build a simple handwritten digit recognition network:

Example

import torch
import torch.nn as nn

class SimpleCNN(nn.Module):
    def __init__(self, num_classes=10):
        super(SimpleCNN, self).__init__()
        # First convolution block
        self.conv1 = nn.Conv2d(1, 32, kernel_size=3, padding=1)
        self.bn1 = nn.BatchNorm2d(32)
        self.relu1 = nn.ReLU()

        # Second convolution block
        self.conv2 = nn.Conv2d(32, 64, kernel_size=3, padding=1)
        self.bn2 = nn.BatchNorm2d(64)
        self.relu2 = nn.ReLU()

        # Pooling layer
        self.pool = nn.MaxPool2d(2, 2)

        # Fully connected layer
        self.fc = nn.Linear(64 * 7 * 7, num_classes)

    def forward(self, x):
        # First convolution block
        x = self.conv1(x)
        x = self.bn1(x)
        x = self.relu1(x)
        x = self.pool(x)

        # Second convolution block
        x = self.conv2(x)
        x = self.bn2(x)
        x = self.relu2(x)
        x = self.pool(x)

        # Flatten and classify
        x = x.view(x.size(0), -1)
        x = self.fc(x)
        return x

# Create the model
model = SimpleCNN(num_classes=10)

# Test input: batch=4, grayscale image 28x28
input_image = torch.randn(4, 1, 28, 28)
output = model(input_image)

print("Input shape:", input_image.shape)
print("Output shape:", output.shape)  # torch.Size([4, 10])

Output Size Calculation

The formula for calculating the output size of a convolution layer:

H_out = floor((H_in + 2 * padding[0] - dilation[0] * (kernel_size[0] - 1) - 1) / stride[0]) + 1
W_out = floor((W_in + 2 * padding[1] - dilation[1] * (kernel_size[1] - 1) - 1) / stride[1]) + 1

Common Questions

Q1: How to choose the kernel size?

Common kernel sizes:

  • 1x1: Used to change the number of channels, add non-linearity
  • 3x3: Most commonly used, balances parameter count and receptive field
  • 5x5、7x7: Larger receptive field, but larger number of parameters

Q2: How to choose padding and stride?

  • To keep the feature map size, usepadding = (kernel_size - 1) / 2
  • For downsampling, usestride > 1

Use Cases

nn.Conv2dIt is one of the most important layers in computer vision. Main application scenarios include:

  • Image Classification: Extract image features, such as VGG, ResNet, etc.
  • Object Detection: YOLO, Faster R-CNN, etc.
  • Semantic Segmentation: U-Net, FCN, etc.
  • Style Transfer: Generate artistic images

Tip: In modern CNNs, 3x3 convolution is the most commonly used. It can cover sufficient spatial information while keeping a relatively small number of parameters.


PyTorch torch.nn 参考手册PyTorch torch.nn Reference Manual

Other Extensions