Common Network Types

Imagine we are teaching a child to recognize cats and dogs.

Initially, we show them many pictures of cats and dogs and tell them "this is a cat, this is a dog." Gradually, the child's brain extracts patterns from these images: cats usually have pointed ears and rounder faces; dogs' ears may droop and their faces are longer. This process is essentiallylearning featuresandand building patterns.。

Deep learning, a powerful branch of machine learning, is centered on enabling computers to simulate this process. By building multi-layer neural networks, it allows machines to automatically learn and extract complex features from massive data, ultimately accomplishing advanced tasks such as image recognition, language understanding, and trend prediction. Different tasks require networks with different architectures. This article will introduce you to several core, common types of deep learning networks and help you understand their design philosophies and typical applications.

Deep Learning Networks

Chinese Full Name English Full Name Abbreviation
Artificial Neural Network Artificial Neural Network ANN
Convolutional Neural Network Convolutional Neural Network CNN
Recurrent Neural Network Recurrent Neural Network RNN
Long Short-Term Memory Network Long Short-Term Memory LSTM
Gated Recurrent Unit Gated Recurrent Unit GRU
Generative Adversarial Network Generative Adversarial Network GAN
Transformer Transformer Transformer
Autoencoder Autoencoder AE
Variational Autoencoder Variational Autoencoder VAE
Deep Belief Network Deep Belief Network DBN
Graph Neural Network Graph Neural Network GNN

Neural Network Basics and Fully Connected Networks

Before diving into various networks, we need to understand the most basic model —fully connected network, also known asmulti-layer perceptron (MLP)。

Core Idea: Everything Can Be Connected

The fully connected network is the most straightforward architecture in deep learning. As the name suggests, each layer'severy neuronis connected to the adjacent layer'severy neuron.

We can think of it as an extremely dense information processing network. Data enters from the input layer, undergoes transformations through multiple hidden layers, and finally produces results at the output layer.

Typical Applications and Limitations

Application: Thanks to its powerful fitting capability, FCN is well suited for structured data (e.g., tabular data such as house area, location, and number of rooms in house price prediction).

Example

# A simple fully connected network example (using PyTorch)
import torch.nn as nn

class SimpleFCN(nn.Module):
    def __init__(self, input_size, num_classes):
        super(SimpleFCN, self).__init__()
        # Define network layers
        self.fc1 = nn.Linear(input_size, 128)  # First hidden layer
        self.relu = nn.ReLU()                  # Activation function
        self.fc2 = nn.Linear(128, 64)         # Second hidden layer
        self.fc3 = nn.Linear(64, num_classes) # Output layer

    def forward(self, x):
        x = self.fc1(x)
        x = self.relu(x)
        x = self.fc2(x)
        x = self.relu(x)
        x = self.fc3(x)
        return x

# Assume input is a 100-dimensional feature, for 10-class classification
model = SimpleFCN(input_size=100, num_classes=10)

Limitations: When dealing with grid-like data such as images and speech, FCN faces significant challenges. Because pixels in an image are highly correlated in space, FCN ignores this spatial structure and flattens the image into a one-dimensional vector for processing, leading to an explosion in the number of parameters and difficulty in learning effective spatial features.


Convolutional Neural Networks — The Cornerstone of Computer Vision

To solve the problem of image processing,Convolutional Neural Networks (CNN)emerged as the times require, completely transforming the field of computer vision.

Core Idea: Local Perception and Parameter Sharing

CNN draws design inspiration from the biological visual cortex, and its two core ideas are:

  1. Local perception: Unlike FCN, where neurons connect to the entire image, each neuron in CNN only perceives asmall local region(e.g., a 3x3 or 5x5 pixel patch). This better matches the property that adjacent pixels in an image are more strongly correlated.
  2. Parameter sharing: Using the sameconvolution kernel(or filter) to slide and scan at different positions of the image, extracting the same type of features (such as edges, textures). This greatly reduces network parameters.

Core Components and Typical Applications

A typical CNN is stacked from the following components:

  • Convolutional layer: Uses convolution kernels to extract features.
  • Pooling layer(e.g., max pooling): Downsamples feature maps, reducing data volume and enhancing feature invariance.
  • Fully connected layer: At the end of the network, maps learned distributed features to the sample label space.

Application: Almost all computer vision tasks such as image classification, object detection, and face recognition.

Example

# A simple CNN example (for image classification)
class SimpleCNN(nn.Module):
    def __init__(self, num_classes=10):
        super(SimpleCNN, self).__init__()
        self.conv1 = nn.Conv2d(in_channels=3, out_channels=16, kernel_size=3, padding=1)
        self.pool = nn.MaxPool2d(kernel_size=2, stride=2)
        self.conv2 = nn.Conv2d(16, 32, 3, padding=1)
        self.fc1 = nn.Linear(32 * 8 * 8, 256) # Assume the feature map size is 8x8 after two pooling operations
        self.fc2 = nn.Linear(256, num_classes)

    def forward(self, x):
        x = self.pool(nn.functional.relu(self.conv1(x))) # Convolution -> Activation -> Pooling
        x = self.pool(nn.functional.relu(self.conv2(x)))
        x = x.view(-1, 32 * 8 * 8) # Flatten the feature map into a one-dimensional vector
        x = nn.functional.relu(self.fc1(x))
        x = self.fc2(x)
        return x

Recurrent Neural Networks — Experts at Processing Sequential Data

For data such as language, speech, and time series that havesequential dependency relationships, we need a network that can remember historical information. This is theRecurrent Neural Network (RNN)。

Core Idea: Introducing a Memory Mechanism

The core of RNN lies in its recurrent structure. When processing the current input, the network combines the current input with thehidden state from the previous time stepto jointly determine the current output and the hidden state passed to the next time step. This is like reading a sentence: understanding the current word requires relying on previously read words.

Variants and Typical Applications

Basic RNN suffers from the long-term dependency problem, making it difficult to learn information in long sequences. This led to two important variants:

  • Long Short-Term Memory network (LSTM): Through an elaborate gating mechanism (input gate, forget gate, output gate), it selectively remembers important information and forgets useless information, effectively solving the long-sequence dependency problem.
  • Gated Recurrent Unit (GRU): A simplified version of LSTM with a more concise structure and higher computational efficiency, performing comparably on many tasks.

Application: Machine translation, text generation, speech recognition, stock price prediction.

Example

# A simple RNN example (for text sentiment classification)
class SimpleRNN(nn.Module):
    def __init__(self, vocab_size, embed_size, hidden_size, num_classes):
        super(SimpleRNN, self).__init__()
        self.embedding = nn.Embedding(vocab_size, embed_size) # Word embedding layer
        self.rnn = nn.RNN(input_size=embed_size, hidden_size=hidden_size, batch_first=True)
        self.fc = nn.Linear(hidden_size, num_classes)

    def forward(self, x):
        # Shape of x: (batch_size, sequence_length)
        x = self.embedding(x) # After embedding: (batch_size, seq_len, embed_size)
        _, h_n = self.rnn(x)  # h_n is the hidden state of the last time step
        out = self.fc(h_n.squeeze(0)) # Use the final state for classification
        return out

Generative Adversarial Networks — From Learning to Creation

If the previous networks are discriminative models (learning to distinguish data), thenGenerative Adversarial Network (GAN)is an outstanding representative of generative models (learning to create data).

Core Idea: Evolution Through Adversarial Games

GAN draws inspiration from game theory. It consists of two adversarial networks:

  • Generator: Like a forger, its goal is to learn the distribution of real data and generate new data realistic enough to pass as genuine.
  • Discriminator: Like an authentication expert, its goal is to accurately distinguish whether input data comes from the real dataset or the generator.

The two improve together through continuous adversarial training: the generator strives to create more realistic data to fool the discriminator, while the discriminator works to improve its ability to distinguish. Eventually, the generator can produce high-quality new data.

Typical Applications

Application: Image generation, image super-resolution, style transfer, data augmentation.

Example

# Schematic pseudocode for GAN's core training loop
for epoch in range(num_epochs):
    # 1. Train the discriminator: maximize its ability to classify real data as real and generated data as fake
    real_data = get_real_data()
    noise = generate_random_noise()
    fake_data = generator(noise).detach() # Note the detach to prevent the generator from being updated

    d_loss_real = criterion(discriminator(real_data), real_labels)
    d_loss_fake = criterion(discriminator(fake_data), fake_labels)
    d_loss = d_loss_real + d_loss_fake
    d_loss.backward()
    optimizer_D.step()

    # 2. Train the generator: minimize the discriminator's ability to classify generated data as fake (i.e., fool the discriminator)
    noise = generate_random_noise()
    fake_data = generator(noise)
    g_loss = criterion(discriminator(fake_data), real_labels) # Make the discriminator believe the generated data is real
    g_loss.backward()
    optimizer_G.step()

Summary and Comparison

Network Type Core Idea Data Types It Excels At Processing Typical Applications
Fully Connected Network Global connection, dense fitting Structured data (tables) House price prediction, credit scoring
Convolutional Neural Network Local perception, parameter sharing Grid data (images) Image classification, object detection
Recurrent Neural Network Temporal dependency, memory state Sequence data (text, time series) Machine translation, speech recognition
Generative Adversarial Network Adversarial game, data generation Used to generate data similar to real data Image generation, style transfer
Other extensions