PyTorch GPU / CUDA Acceleration

The core operations of deep learning are large-scale matrix multiplication and element-wise operations. CPUs are designed to handle complex serial logic, typically with 8 to 64 cores; GPUs, on the other hand, have thousands of simple parallel cores, making them naturally suited for this kind of highly parallel numerical computation. PyTorch, through NVIDIA'sCUDA(Compute Unified Device Architecture) framework, invokes the GPU, which can speed up training by dozens or even a hundred times.


1. Differences between CPU and GPU

The improvement in GPU training speed mainly comes from two aspects: first, the parallel execution of a large number of identical computations; second, high-bandwidth GPU memory enables data transfer much faster than CPU memory.

For compute-intensive operations such as matrix multiplication, the acceleration effect is particularly significant.

Comparison item CPU GPU(NVIDIA)
Number of cores 8-64 large cores Thousands of small cores (CUDA Cores)
Design goal Low-latency serial processing High-throughput parallel computing
Memory bandwidth ~50~100 GB/s ~500~3000 GB/s
Matrix multiplication speed Baseline 10x-100x faster
PyTorch interface "cpu" "cuda"

If using Apple Silicon (M1/M2/M3), PyTorch, through thempsbackend, supports Metal GPU acceleration. The usage is almost identical to CUDA; just change the device to"mps"。


2. Detecting the CUDA Environment

Before using the GPU, you need to confirm whether the current environment has a CUDA-enabled PyTorch installed.

Also, check whether there are any available GPU devices on the system.

Example

import torch

# Whether CUDA is supported
print("CUDA available:", torch.cuda.is_available())

# Number of GPU devices
print("GPU count:", torch.cuda.device_count())

# Index of the current default GPU
print("Current GPU:", torch.cuda.current_device())

# GPU model name
print("GPU model:", torch.cuda.get_device_name(0))

# PyTorch version and the CUDA version it was compiled with
print("PyTorch version:", torch.__version__)
print("CUDA version:", torch.version.cuda)

Example output:

CUDA 可用: True
GPU 数量: 1
当前 GPU: 0
GPU 型号: NVIDIA GeForce RTX 4090
PyTorch 版本: 2.3.0+cu121
CUDA 版本: 12.1

2.1 Dynamically Selecting a Device (Recommended Approach)

Hardcoding in code"cuda"This will cause machines without a GPU to directly throw an error.

Example

import torch

# Method 1: Classic approach, best compatibility
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# Method 2: Recommended for PyTorch 2.0+, automatically supports CUDA / MPS / CPU
device = (
    "cuda" if torch.cuda.is_available()
    else "mps" if torch.backends.mps.is_available()
    else "cpu"
)

print(f"Using device: {device}")

3. Moving Tensors Between Devices

Tensors in PyTorch are created on the CPU by default.

To use GPU computation, you need to explicitly move tensors to the GPU, or create them directly on the GPU.

3.1 Basic Moving Methods

Example

import torch

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# Create a CPU tensor
cpu_tensor = torch.tensor([1.0, 2.0, 3.0])
print(cpu_tensor.device)   # cpu

# Method 1: .to(device) — recommended, most versatile
gpu_tensor = cpu_tensor.to(device)

# Method 2: .cuda() — CUDA environment only
gpu_tensor = cpu_tensor.cuda()

# Method 3: Specify the device directly when creating
gpu_tensor = torch.tensor([1.0, 2.0, 3.0], device=device)
gpu_tensor = torch.randn(3, 4, device=device)

# Move back to CPU (for printing, numpy conversion, saving, etc.)
back_to_cpu = gpu_tensor.cpu()

print(gpu_tensor.device)    # cuda:0
print(back_to_cpu.device)   # cpu

When converting a GPU tensor to numpy, you must first move it back to the CPU, and if the tensor has gradients, you also need to detach it first:

Example

# Convert an ordinary GPU tensor to numpy
arr = gpu_tensor.cpu().numpy()

# Convert a GPU tensor with gradients to numpy
arr = gpu_tensor.detach().cpu().numpy()

3.2 Device Consistency Constraint

Tensors on different devices cannot directly participate in the same operation.

Otherwise, it will throwRuntimeError:

Example

a = torch.randn(3).to("cuda")
b = torch.randn(3)           # On CPU

# c = a + b    # RuntimeError: Expected all tensors to be on the same device

# Correct approach: unify the device first
c = a + b.to("cuda")

Check the device where the tensor is located:

Example

x = torch.randn(3, 4).to(device)

print(x.device)           # cuda:0
print(x.is_cuda)          # True
print(x.get_device())     # 0 (GPU index)

3.3 Speed Comparison Verification

Example

import torch
import time

device = torch.device("cuda")
n = 5000

# CPU matrix multiplication
a_cpu = torch.randn(n, n)
b_cpu = torch.randn(n, n)
start = time.time()
c_cpu = torch.matmul(a_cpu, b_cpu)
print(f"CPU time: {time.time() - start:.3f}s")

# GPU matrix multiplication
a_gpu = a_cpu.to(device)
b_gpu = b_cpu.to(device)
torch.cuda.synchronize()    # Make sure data transfer is complete before starting the timer
start = time.time()
c_gpu = torch.matmul(a_gpu, b_gpu)
torch.cuda.synchronize()    # Wait for the GPU to finish execution before stopping the timer
print(f"GPU time: {time.time() - start:.3f}s")

Example output:

CPU 耗时: 1.847s
GPU 耗时: 0.021s

GPU computation isasynchronously executed—after the Python call returns, the GPU operation may not be complete. When timing, you must calltorch.cuda.synchronize()to wait for the GPU to actually finish; otherwise, the measurement results are inaccurate.


4. Moving the Model to GPU

All parameters of the model (weight、bias) are essentially tensors.

These parameters also need to be moved to the GPU in order to perform forward and backward propagation on the GPU.

Call.to(device)on the entire model, and PyTorch will automatically traverse and move all internal parameters.

Example

import torch
import torch.nn as nn

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

class SimpleNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(784, 256),
            nn.ReLU(),
            nn.Linear(256, 128),
            nn.ReLU(),
            nn.Linear(128, 10),
        )

    def forward(self, x):
        return self.net(x)

# Move the model to GPU, only need to call once
model = SimpleNet().to(device)

# Verify that all parameters are on the GPU
for name, param in model.named_parameters():
    print(f"{name}: {param.device}")
# net.0.weight: cuda:0
# net.0.bias:   cuda:0
# ...

# Input data must also be on the same device
x = torch.randn(32, 784).to(device)
output = model(x)
print(output.shape)   # torch.Size([32, 10])

When the model is on the GPU but the input data is still on the CPU, the forward pass will raise an error. Be sure to, after DataLoader reads each batch,inputsandlabelscall .to(device) on all of them..to(device)。


5. Complete Training Workflow

The following is a standard GPU training template that includes data loading, model training, and validation evaluation.

Example

import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

# 1. Device configuration
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Training device: {device}")

# 2. Data loading
# pin_memory=True: Locks data in memory to speed up CPU -> GPU transfer
# num_workers: Multiprocess preloading to reduce data waiting time
transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.5,), (0.5,)),
])
train_dataset = datasets.MNIST(root="./data", train=True,
                                download=True, transform=transform)
test_dataset  = datasets.MNIST(root="./data", train=False,
                                download=True, transform=transform)

train_loader = DataLoader(train_dataset, batch_size=256, shuffle=True,
                          num_workers=4, pin_memory=True)
test_loader  = DataLoader(test_dataset,  batch_size=256, shuffle=False,
                          num_workers=4, pin_memory=True)

# 3. Define the model
class CNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(1, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
            nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
        )
        self.classifier = nn.Sequential(
            nn.Flatten(),
            nn.Linear(64 * 7 * 7, 256), nn.ReLU(), nn.Dropout(0.5),
            nn.Linear(256, 10),
        )

    def forward(self, x):
        return self.classifier(self.features(x))

model     = CNN().to(device)     # Move the model to GPU
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=1e-3)

# 4. Training function
def train_epoch(model, loader, optimizer, criterion):
    model.train()
    total_loss, correct = 0.0, 0
    for inputs, labels in loader:
        # non_blocking=True: Asynchronous transfer, CPU can continue preparing the next batch of data
        inputs = inputs.to(device, non_blocking=True)
        labels = labels.to(device, non_blocking=True)

        optimizer.zero_grad()
        outputs = model(inputs)
        loss    = criterion(outputs, labels)
        loss.backward()
        optimizer.step()

        total_loss += loss.item() * inputs.size(0)
        correct    += (outputs.argmax(1) == labels).sum().item()

    n = len(loader.dataset)
    return total_loss / n, correct / n

# 5. Validation function
def eval_epoch(model, loader, criterion):
    model.eval()
    total_loss, correct = 0.0, 0
    with torch.no_grad():
        for inputs, labels in loader:
            inputs = inputs.to(device, non_blocking=True)
            labels = labels.to(device, non_blocking=True)
            outputs = model(inputs)
            loss    = criterion(outputs, labels)
            total_loss += loss.item() * inputs.size(0)
            correct    += (outputs.argmax(1) == labels).sum().item()
    n = len(loader.dataset)
    return total_loss / n, correct / n

# 6. Main training loop
for epoch in range(1, 11):
    train_loss, train_acc = train_epoch(model, train_loader, optimizer, criterion)
    val_loss,   val_acc   = eval_epoch(model,  test_loader,  criterion)
    print(f"Epoch {epoch:02d} | "
          f"Train Loss: {train_loss:.4f}, Acc: {train_acc:.4f} | "
          f"Val Loss: {val_loss:.4f}, Acc: {val_acc:.4f}")

Key points:

  • deviceVariables are declared at the very top of the code, referenced globally, never hardcoded"cuda"
  • Model.to(device)only needs to be called once
  • For each batch,inputsandlabelsmust also be.to(device)
  • DataLoader settingspin_memory=Trueandnon_blocking=Trueaccelerate data transfer
  • Use in validation phasetorch.no_grad()disable gradients to save memory

6. Multi-GPU Training

When a single GPU's memory is insufficient, consider using multi-GPU parallel training.

In addition, multi-GPU can further improve training speed. PyTorch provides two main approaches.

6.1 DataParallel

This is the simplest multi-GPU approach.

It uses a single process, splits each batch evenly across GPUs, runs forward propagation in parallel, and aggregates gradients on the main GPU for updates. It is easy to use, but the main GPU is heavily loaded, utilization across GPUs is uneven, and it is suitable for quick starts.

Example

import torch
import torch.nn as nn

model = CNN()

if torch.cuda.device_count() > 1:
    print(f"Using {torch.cuda.device_count()} GPUs")
    model = nn.DataParallel(model)

    # You can also specify which GPUs to use
    # model = nn.DataParallel(model, device_ids=[0, 1])

model = model.to("cuda")

# The subsequent training code is exactly the same as single-GPU
# DataParallel automatically splits the batch evenly across GPUs and aggregates the results

If you need to access the original model's attributes (such as custom methods), you need to.moduleaccess them via:

Example

# After model is wrapped by DataParallel, the original model is in model.module
print(model.module.classifier)

# When saving the model, it is recommended to save model.module for easy single-GPU loading
torch.save(model.module.state_dict(), "model.pth")

6.2 DistributedDataParallel

This is the multi-process approach.

Each process is bound to one GPU, each card holds a full model copy, and gradients are synchronized via All-Reduce. It is the recommended solution for production environments and large-scale training, with far higher efficiency than DataParallel.

Example

import torch
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.utils.data.distributed import DistributedSampler

def main(rank, world_size):
    # Initialize process group; the NCCL backend is optimized for NVIDIA GPUs
    dist.init_process_group(backend="nccl", rank=rank, world_size=world_size)

    # Bind each process to its corresponding GPU
    torch.cuda.set_device(rank)
    device = torch.device(f"cuda:{rank}")

    # Wrap with DDP after moving the model to the corresponding GPU
    model = CNN().to(device)
    model = DDP(model, device_ids=[rank])

    # Use DistributedSampler in DataLoader to ensure no data overlap across GPUs
    sampler = DistributedSampler(train_dataset,
                                  num_replicas=world_size, rank=rank)
    loader  = DataLoader(train_dataset, batch_size=64,
                         sampler=sampler, pin_memory=True)

    # Training logic is identical to single-GPU
    # ...

    dist.destroy_process_group()

# How to launch (recommended: torchrun):
# torchrun --nproc_per_node=4 train_ddp.py
Comparison item DataParallel DistributedDataParallel
Number of processes Single process Multi-process (one per GPU)
Communication backend Python GIL limitation NCCL (efficient)
Main GPU load Heavy (gradient aggregation) Balanced (All-Reduce)
Code changes Minimal Moderate
Applicable scenarios Rapid experimentation Production training

7. Mixed Precision Training AMP

Training uses float32 (FP32) precision by default.

Automatic Mixed Precision (AMP)Uses float16 (FP16) or bfloat16 (BF16) for some computations, yielding significant benefits with almost no loss of precision:

  • Memory usage reduced by about 50%
  • Training speed improved 2x–3x (Tensor Core hardware acceleration)
  • Very few code changes required; only three modifications

Example

from torch.cuda.amp import autocast, GradScaler

model     = CNN().to(device)
optimizer = optim.Adam(model.parameters(), lr=1e-3)
scaler    = GradScaler()    # FP16 gradient scaler to prevent gradients from underflowing to zero

for epoch in range(num_epochs):
    model.train()
    for inputs, labels in train_loader:
        inputs = inputs.to(device, non_blocking=True)
        labels = labels.to(device, non_blocking=True)

        optimizer.zero_grad()

        # Modification 1: automatically select FP16/FP32 within the autocast region
        with autocast(device_type="cuda"):
            outputs = model(inputs)
            loss    = criterion(outputs, labels)

        # Modification 2: scale the loss with scaler before backpropagation
        scaler.scale(loss).backward()

        # Modification 3: scaler updates parameters and automatically handles gradient scaling internally
        scaler.step(optimizer)
        scaler.update()

7.1 Choosing Between FP16 and BF16

Format Precision bits Exponent bits Suitable hardware Characteristics
float16 10 bits 5 bits RTX 20/30/40 series Requires GradScaler to prevent overflow
bfloat16 7 bits 8 bits A100 / H100 / RTX 4090 Same range as FP32, more stable training

If your GPU supports BF16, use it preferentially; no GradScaler needed:

Example

# BF16 usage (PyTorch 2.0+)
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
    outputs = model(inputs)
    loss    = criterion(outputs, labels)

loss.backward()
optimizer.step()

8. Performance Optimization Tips

This section introduces several techniques for optimizing memory usage and training speed.

8.1 GPU Memory Management

Example

# View current memory usage
print(f"Allocated: {torch.cuda.memory_allocated() / 1024**2:.1f} MB")
print(f"Cached: {torch.cuda.memory_reserved() / 1024**2:.1f} MB")

# Print detailed memory report
print(torch.cuda.memory_summary())

# Empty the cache pool (does not release allocated memory)
torch.cuda.empty_cache()

# Use inference_mode during inference (faster than no_grad, completely disables the gradient engine)
with torch.inference_mode():
    output = model(x)

# Gradient checkpointing: trade recomputation for memory (commonly used in large model training)
from torch.utils.checkpoint import checkpoint
output = checkpoint(model_block, x)

8.2 DataLoader Optimization

Example

loader = DataLoader(
    dataset,
    batch_size=256,
    num_workers=4,              # Multi-process prefetching; recommended to set to 50% of CPU cores
    pin_memory=True,            # Pinned memory to speed up CPU -> GPU transfer
    persistent_workers=True,    # Keep worker processes alive to avoid recreating processes every epoch
    prefetch_factor=2,          # Number of batches prefetched per worker
)

8.3 torch.compile Model Compilation (PyTorch 2.0+)

One line of code compiles and optimizes the computation graph, improving speed by 30%–200% (depending on model architecture):

Example

# Compile the model; the first run has compilation overhead, subsequent iterations are significantly faster
model = torch.compile(model)

# Trade-offs between different compilation modes
model = torch.compile(model, mode="default")          # Balanced, suitable for most scenarios
model = torch.compile(model, mode="reduce-overhead")  # Reduces scheduling overhead
model = torch.compile(model, mode="max-autotune")     # Maximum optimization, longer compilation time

8.4 Gradient Accumulation (Simulating Large Batches)

When memory is insufficient, gradient accumulation can simulate a larger batch size without actually increasing memory usage:

Example

accumulation_steps = 8    # Update parameters every 8 batches, equivalent to batch_size × 8

for i, (inputs, labels) in enumerate(train_loader):
    inputs = inputs.to(device)
    labels = labels.to(device)

    outputs = model(inputs)
    loss    = criterion(outputs, labels) / accumulation_steps   # Divide loss equally
    loss.backward()    # Accumulate gradients without zeroing

    if (i + 1) % accumulation_steps == 0:
        optimizer.step()
        optimizer.zero_grad()    # Only zero gradients after parameter update

8.5 Strategies for Insufficient GPU Memory

Strategy Method Effect
Reduce batch size 256 -> 64 Linearly reduces memory
Mixed precision AMP + FP16/BF16 Memory reduced by about 50%
Gradient accumulation Update every N steps Simulate large batch without increasing memory
Gradient checkpointing torch.utils.checkpoint Greatly reduces memory, but reduces training speed
Freeze some parameters Freeze backbone in transfer learning Reduces memory usage of backpropagation

9. Common Errors and Troubleshooting

This section summarizes the most common errors in GPU training and their solutions.

9.1 RuntimeError: Expected all tensors to be on the same device

Cause: The tensors involved in the operation are on different devices (one on CPU, one on GPU).

Example

# Troubleshooting: print the device of each tensor
print(inputs.device, labels.device, next(model.parameters()).device)

# Solution: ensure .to(device) is executed for every batch
inputs = inputs.to(device)
labels = labels.to(device)

9.2 RuntimeError: CUDA out of memory

Cause: Out of memory, commonly due to batch size being too large, model being too large, or accidental accumulation of computation graphs.

Example

# Common pitfall: loss.item() is not called in the loop, causing computation graphs to accumulate
# Incorrect code
total_loss += loss           # loss is a tensor that holds the entire computation graph

# Correct code
total_loss += loss.item()    # .item() extracts a Python scalar and releases the computation graph

# Other troubleshooting steps:
# 1. Reduce batch_size
# 2. Ensure the validation loop uses torch.no_grad()
# 3. Call torch.cuda.empty_cache() to clear cache
# 4. Use torch.cuda.memory_summary() to locate the source of GPU memory usage

9.3 Can't call numpy() on Tensor that requires grad

Reason: Tensors with gradients cannot be directly converted to numpy arrays.

Example

# Incorrect code
arr = gpu_tensor.numpy()

# Correct code
arr = gpu_tensor.detach().cpu().numpy()

9.4 CUDA error: device-side assert triggered

Reason: Usually it is label values out of bounds (e.g., 10 classes but the label value equals 10), or array index out of bounds. The error message is generated on the GPU, and the default display location is inaccurate.

Example

# Set environment variable so errors are thrown at the exact code line (becomes synchronous execution, slower)
import os
os.environ["CUDA_LAUNCH_BLOCKING"] = "1"

9.5 Training Speed Not Improving (Low GPU Utilization)

Reason: Data loading becomes the bottleneck; the GPU spends most of its time waiting for the CPU to prepare data.

Troubleshooting steps:

  1. Runnvidia-smiorwatch -n 1 nvidia-smiObserve GPU utilization
    • Utilization consistently below 80% indicates a data bottleneck
  2. Increase DataLoader's num_workers
  3. Enable pin_memory=True
  4. Move data preprocessing to the GPU (torchvision.transforms supports GPU operations)
  5. Preload small-file datasets into memory

9.6 Using PyTorch Profiler to Pinpoint Bottlenecks

Example

from torch.profiler import profile, ProfilerActivity

with profile(
    activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
    record_shapes=True,
) as prof:
    for i, (inputs, labels) in enumerate(train_loader):
        if i >= 10:
            break
        inputs = inputs.to(device)
        labels = labels.to(device)
        loss   = criterion(model(inputs), labels)
        loss.backward()
        optimizer.step()
        optimizer.zero_grad()

# Sort by CUDA time, print top 15 items
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=15))

10. API Quick Reference

The following is a quick reference table for common PyTorch GPU-related APIs.

10.1 Device Management

Operation Code
Detect CUDA availability torch.cuda.is_available()
Select device torch.device("cuda" if ... else "cpu")
Number of GPUs torch.cuda.device_count()
GPU name torch.cuda.get_device_name(0)
Set current GPU torch.cuda.set_device(0)
Wait for GPU to complete torch.cuda.synchronize()

10.2 Tensor Operations

Operation Code
Move tensor to GPU tensor.to(device)ortensor.cuda()
Move tensor to CPU tensor.cpu()
View current device tensor.device
Whether it is on GPU tensor.is_cuda
Convert GPU tensor to numpy tensor.detach().cpu().numpy()
Asynchronous transfer tensor.to(device, non_blocking=True)

10.3 Model Operations

Operation Code
Move model to GPU model.to(device)
Enable training mode model.train()
Enable inference mode model.eval()
Compile acceleration torch.compile(model)
Disable gradients with torch.no_grad():
Inference mode with torch.inference_mode():

10.4 GPU Memory Management

Operation Code
Allocated GPU memory torch.cuda.memory_allocated()
Cached GPU memory torch.cuda.memory_reserved()
GPU memory report torch.cuda.memory_summary()
Clear cache torch.cuda.empty_cache()

10.5 Core Principles Quick Reference

1. 用 device 变量统一管理,不要硬编码 "cuda"
2. 模型和数据必须在同一设备上,每个 batch 都要 .to(device)
3. 验证和推理时一定使用 torch.no_grad() 或 torch.inference_mode()
4. 生产训练推荐开启 AMP 混合精度,几乎免费获得 2x 加速
5. DataLoader 设置 pin_memory=True 和 num_workers >= 4 减少数据瓶颈
6. PyTorch 2.0+ 可以用 torch.compile(model) 一行提速 30%+
7. GPU 计算是异步的,精确计时需要调用 torch.cuda.synchronize()
Other extensions