PyTorch Introduction
PyTorch is an open-source Python machine learning library based on the Torch library, implemented at the underlying level in C++, and applied in artificial intelligence fields such as computer vision and natural language processing.
PyTorch was originally developed by Meta Platforms' AI research team and is now part of the Linux Foundation.
Many deep learning software packages are built on PyTorch, including Tesla Autopilot, Uber's Pyro, Hugging Face's Transformers, PyTorch Lightning, and Catalyst.
PyTorch has two main features:
- Tensor computation similar to NumPy, which can be accelerated on hardware accelerators such as GPU or MPS.
- Deep neural networks based on an automatic differentiation system.
PyTorch includes submodules such as torch.autograd, torch.nn, and torch.optim.
PyTorch includes a variety of loss functions, including MSE (mean squared error = L2 norm), cross-entropy loss, and negative log-likelihood loss (useful for classifiers), among others.

PyTorch Features
-
Dynamic Computation Graphs: PyTorch's computation graphs are dynamic, meaning they are built at runtime and can be changed at any time. This provides great flexibility for experimentation and debugging, because developers can execute code line by line and inspect intermediate results.
-
Automatic Differentiation: PyTorch's automatic differentiation system allows developers to easily compute gradients, which is crucial for training deep learning models. It automatically calculates the gradients of the loss function with respect to model parameters through the backpropagation algorithm.
-
Tensor Computation: PyTorch provides tensor operations similar to NumPy, which can be executed on CPU and GPU, thereby accelerating the computation process. Tensors are the basic data structure in PyTorch, used to store and manipulate data.
-
Rich API: PyTorch provides a large number of predefined layers, loss functions, and optimization algorithms, all of which are common components for building deep learning models.
-
Multi-language Support: Although PyTorch uses Python as its main interface, it also provides a C++ interface, allowing lower-level integration and control.
Dynamic Computation Graph
One of the most notable features of PyTorch is its dynamic computation graph mechanism.
Unlike TensorFlow's static computation graph, PyTorch builds the computation graph during execution, meaning that with each computation, the graph automatically changes according to the shape of the input data.
Advantages of the dynamic computation graph:
- More flexible, especially suitable for scenarios requiring conditional judgment or recursion.
- Convenient for debugging and modification, allowing direct inspection of intermediate results.
- Closer to Python programming style and easy to get started.
Tensor and Autograd
The core data structure in PyTorch is theTensor, which is a multi-dimensional matrix that can efficiently perform computations on CPU or GPU. Tensor operations support the Autograd mechanism, enabling automatic gradient computation during backpropagation, which is crucial for gradient descent optimization algorithms in deep learning.
Tensor:
- Supports switching between CPU and GPU.
- Provides a NumPy-like interface and supports element-wise operations.
- Supports automatic differentiation, making gradient computation convenient.
Autograd:
- PyTorch's built-in automatic differentiation engine can automatically track all tensor operations and compute gradients during backpropagation.
- Through
requires_gradattribute, you can specify that a tensor requires gradient computation. - Supports efficient backpropagation, suitable for neural network training.
Model Definition and Training
PyTorch providestorch.nnmodule, allowing users to inheritnn.Moduleclass to define neural network models. Useforwardfunction to specify forward propagation; automatic backpropagation (throughautograd) and gradient computation are also handled internally by PyTorch.
Neural network module (torch.nn):
- Provides commonly used layers (such as linear layers, convolutional layers, pooling layers, etc.).
- Supports defining complex neural network architectures (including networks with multiple inputs and outputs).
- Compatible with optimizers (such as
torch.optim) for use together.
GPU Acceleration
PyTorch fully supports running on GPU to accelerate the training of deep learning models. Through a simple.to(device)method, users can transfer models and tensors to the GPU for computation. PyTorch supports multi-GPU training and can leverage NVIDIA CUDA technology to significantly improve computational efficiency.
GPU support:
- Automatically selects GPU or CPU.
- Supports acceleration via CUDA.
- Supports multi-GPU parallel computing (
DataParallelortorch.distributed)。
Ecosystem and Community Support
As an open-source project, PyTorch has a large community and ecosystem. It is not only widely used in academia, but also widely deployed in industry, especially in fields such as computer vision and natural language processing. PyTorch also provides many tools and libraries related to deep learning, such as:
- torchvision: datasets and models for computer vision tasks.
- torchtext: datasets and preprocessing tools for natural language processing tasks.
- torchaudio: a toolkit for audio processing.
- PyTorch Lightning: a high-level library that simplifies PyTorch code, focusing on rapid iteration in research and experimentation.
Comparison with Other Frameworks
Due to its flexibility, ease of use, and community support, PyTorch has become the preferred framework for many deep learning researchers and developers.
TensorFlow vs PyTorch
- PyTorch's dynamic computation graph makes it more flexible and suitable for rapid experimentation and research, while TensorFlow's static computation graph has more room for optimization in production environments.
- PyTorch is more convenient for debugging, while TensorFlow is more mature in deployment and supports a wider range of hardware and platforms.
- In recent years, TensorFlow has also introduced dynamic graphs (such as TensorFlow 2.x), making the two increasingly close in functionality.
- Other deep learning frameworks, such as Keras and Caffe, also have certain applications, but due to its flexibility, ease of use, and community support, PyTorch has become the preferred framework for many deep learning researchers and developers.
| Feature | TensorFlow | PyTorch |
|---|---|---|
| Developer | Facebook (FAIR) | |
| Computation graph type | Static computation graph (defined before execution) | Dynamic computation graph (executed upon definition) |
| Flexibility | Low (computation graph is built at compile time and not easy to modify) | High (computation graph is dynamically created at execution time, easy to modify and debug) |
| Debugging | Difficult (requires usingtf.debuggingor external tools for debugging) | Easy (can directly debug in Python) |
| Ease of use | Low (more complex, more APIs, steeper learning curve) | High (concise API, syntax closer to Python, easy to get started) |
| Deployment | Strong (supports a wide range of hardware, such as TensorFlow Lite and TensorFlow.js) | Weaker (relatively few deployment tools and platforms, although there is TensorFlow support) |
| Community support | Very strong (mature and large community, extensive tutorials and documentation) | Very strong (active community, especially in academia, rapidly developing ecosystem) |
| Model training | Supports distributed training and supports multiple devices (such as CPU, GPU, TPU) | Supports distributed training, supports multi-GPU, CPU, and TPU |
| API level | High-level API: Keras; low-level API: TensorFlow Core | High-level APIs: TorchVision, TorchText, etc.; low-level API: Torch |
| Performance | High (mature in optimization, suitable for production environments) | High (suitable for research and prototyping, production performance is also improving) |
| Automatic differentiation | Supportedtf.GradientTapeImplements dynamic differentiation (more complex) | SupportedautogradDynamic differentiation (more concise and intuitive) |
| Tuning and scalability | Strong (supports running on multiple platforms, such as TensorFlow Serving, etc.) | Weaker (although it excels in academic and experimental environments, production environment support is relatively limited) |
| Framework flexibility | Lower (TensorFlow 2.x introduced dynamic graph features, but it is still not fully flexible) | High (dynamic graph brings higher flexibility) |
| Supports multiple languages | Supports multiple languages (Python, C++, Java, JavaScript, etc.) | Mainly supports Python (but also has a C++ API) |
| Compatibility and migration | TensorFlow 2.x has good compatibility with older versions | Poor compatibility with TensorFlow, migration is difficult |
PyTorch vs NumPy
| Features | PyTorch | NumPy |
|---|---|---|
| Goal | Dedicated to deep learning | General scientific computing |
| GPU support | Natively supports CUDA | Not directly supported |
| Automatic differentiation | Built-in automatic differentiation | Requires manual gradient computation |
| Neural networks | Rich set of neural network modules | Need to implement from scratch |
| Learning cost | Relatively high | Relatively low |
History and Development of PyTorch
PyTorch's predecessor was Torch, a scientific computing framework based on the Lua language. With the growing popularity of Python in machine learning, the Facebook team decided to port Torch's core ideas to Python, giving birth to PyTorch.
- 2016: Facebook released PyTorch version 0.1
- 2017: PyTorch 0.2 introduced distributed training support
- 2018: PyTorch 1.0 was released, adding production deployment capabilities
- 2019: PyTorch 1.3 introduced mobile support
- 2020: PyTorch 1.6 added automatic mixed-precision training
- 2021: PyTorch 1.9 introduced TorchScript and C++ frontend
- 2022: PyTorch 1.12 optimized performance and stability
- 2023: PyTorch 2.0 was released, introducing a compilation mode that greatly improves performance