BERT Series Models

BERT (Bidirectional Encoder Representations from Transformers) is a revolutionary natural language processing model proposed by Google in 2018, which completely changed the research and application paradigms in the NLP field.

This article will systematically introduce BERT's core principles, training methods, fine-tuning techniques, and mainstream variant models.


BERT Architecture and Training

The figure below showsBERT(Bidirectional Encoder Representations from Transformers)the model's core architecture and the masked language modeling (Masked Language Modeling, MLM) task during pretraining.

1. Input Layer (Embedding)

  • Input sequence: text composed of words (or subwords), for example[W₁, W₂, W₃, [MASK], W₅, W₆, W₇, W₂, W₃, W₄, W₅]。
    • [MASK]is the word randomly masked by BERT during pretraining (as in the original text,W₄is replaced by[MASK])。
  • Embedding layer: converts each word into a fixed-dimensional vector representation (e.g., 768 dimensions), including:
    • Token embeddings: semantic information of the vocabulary.
    • Position embeddings: position information of words in the sequence.
    • Segment embeddings: distinguish sentences (useful for sentence-pair tasks, not explicitly shown in the figure).

2. Transformer Encoder

  • Multi-layer Transformer blocks: details not expanded in the figure, but each block contains:
    • Self-attention mechanism: bidirectionally captures contextual dependencies (BERT's core feature).
    • Feed-forward network: nonlinear transformation.
    • Residual connections and layer normalization: stabilize the training process.
  • Output: the context-aware vector representation corresponding to each input word (e.g.,O₁, O₂, ..., O₅)。

3. Masked Language Modeling (MLM) Task

  • Objective: predict the masked word[MASK]corresponding to the original word (in the figureW₄)。
  • Classification layer:
    • Fully-connected layer: maps the Transformer output vector (e.g.,O₄) to the dimension of the vocabulary size.
    • Activation function GELU: Gaussian Error Linear Unit (the nonlinear function used by BERT).
    • Layer normalization (Norm): standardizes the output.
    • Softmax: computes the probability of each word in the vocabulary and selects the word with the highest probability as the prediction result (e.g.,W'₁, W'₂, ..., W'₅is a candidate word).

Transformer Encoder Structure

BERT is built upon the encoder part of the Transformer, whose core is the multi-layer self-attention mechanism:

Example

# Simplified Transformer encoder layer
class TransformerEncoderLayer(nn.Module):
    def __init__(self, d_model, nhead, dim_feedforward=2048):
        super().__init__()
        self.self_attn = MultiheadAttention(d_model, nhead)
        self.linear1 = nn.Linear(d_model, dim_feedforward)
        self.linear2 = nn.Linear(dim_feedforward, d_model)
        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)
       
    def forward(self, src):
        # Self-attention mechanism
        src2 = self.self_attn(src, src, src)[0]
        src = src + self.norm1(src2)
        # Feed-forward network
        src2 = self.linear2(F.relu(self.linear1(src)))
        src = src + self.norm2(src2)
        return src

Key Innovation: Bidirectional Context Modeling

Unlike traditional language models, BERT achieves bidirectional context understanding through the following two pretraining tasks:

  1. Masked Language Model (MLM): randomly masks 15% of input tokens and predicts the masked words
  2. Next Sentence Prediction (NSP): determines whether two sentences appear consecutively

Training Parameters and Configuration

Parameter BERT-base BERT-large
Number of layers 12 24
Hidden layer size 768 1024
Number of attention heads 12 16
Total parameter count 110M 340M

BERT Fine-tuning Methods

Standard Fine-tuning Process

  1. Adding task-specific adaptation layer: add a classification/regression layer according to the downstream task
  2. Learning rate setting: usually use a small learning rate (2e-5 to 5e-5)
  3. Batch size: 16 or 32 are common choices
  4. Training epochs: 2-4 epochs are usually sufficient

Efficient Fine-tuning Techniques

Example

# Example of fine-tuning using HuggingFace Transformers
from transformers import BertForSequenceClassification, Trainer

model = BertForSequenceClassification.from_pretrained('bert-base-uncased')
trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset,
    eval_dataset=eval_dataset
)
trainer.train()

Comparison of Common Fine-tuning Strategies

Method Advantage Disadvantage
Full-parameter fine-tuning Best performance High computational cost
Feature extraction (freeze BERT) Computationally efficient Suboptimal performance
Adapter Parameter-efficient Requires architecture modification
Prompt learning Good for few-shot performance Requires designing prompt templates

Mainstream BERT Variant Models

RoBERTa (Robustly Optimized BERT)

  • Improvement points:
    • Larger batch size (8k vs 256)
    • Longer training time
    • Removed NSP task
    • Dynamic masking pattern
  • Performance: average improvement of 2-3% on the GLUE benchmark

ALBERT (A Lite BERT)

  • Core innovations:
    • Parameter sharing (sharing attention parameters across layers)
    • Embedding factorization (decomposing the word embedding into two small matrices)
  • Effect: parameter count reduced by 89%, speed increased by 1.7 times

Other Important Variants

  1. DistilBERT: compresses the model through knowledge distillation
  2. ELECTRA: replaces MLM with a generator-discriminator architecture
  3. SpanBERT: optimizes modeling of text spans

Chinese BERT Models

Overview of Chinese Pretrained Models

Model Institution Features
BERT-wwm Harbin Institute of Technology Whole Word Masking
RoBERTa-wwm-ext Harbin Institute of Technology Expanded training data
ERNIE (Baidu) Baidu Incorporates knowledge graphs
NEZHA Huawei Relative position encoding

Chinese BERT Usage Example

Example

from transformers import BertTokenizer, BertModel

tokenizer = BertTokenizer.from_pretrained('bert-base-chinese')
model = BertModel.from_pretrained('bert-base-chinese')

inputs = tokenizer("Natural language processing is very interesting", return_tensors="pt")
outputs = model(**inputs)

Fine-tuning Recommendations for Chinese Tasks

  1. Using the Whole Word Masking (wwm) version yields better results
  2. Pay attention to handling Chinese word segmentation boundary issues
  3. For specialized domains, consider domain-adaptive pretraining

Practical Recommendations and Resources

Learning Roadmap

Recommended Resources

  1. Papers:
    • Original BERT paper (arXiv:1810.04805)
    • Papers on variants such as RoBERTa, ALBERT, etc.
  2. Code repositories:
    • HuggingFace Transformers
    • GitHub implementation of Chinese BERT
  3. Online courses:
    • Coursera Natural Language Processing Specialization
    • Hung-yi Lee's Deep Learning Course

Through systematic learning and practice, BERT series models can become a powerful tool for solving NLP problems. It is recommended to start with the basic version and gradually explore more advanced variants and optimization techniques.

Other Extensions