Hugging Face Transformers

Hugging Face Transformers is currently the most popular open-source NLP / AI library, providing thousands of pre-trained models that cover almost all AI tasks including text, image, audio, and multimodal.

Its core value: encapsulating complex model loading, inference, and training processes into a few lines of code.

Hugging Face Ecosystem Overview Hub 400K+ Models 100K+ Datasets Transformers Pretrained Model Inference and Fine-tuning Framework Datasets Massive Datasets Efficient Loading and Processing PEFT Parameter-Efficient Fine-Tuning LoRA / QLoRA Accelerate Multi-GPU/TPU Training Distributed Acceleration Tokenizers High-Performance Tokenizer (Rust implementation) Evaluate Model Evaluation Metrics (BLEU/F1, etc.)

Supported Task Types

Task Categories Supported by Transformers NLP Natural Language Processing Text Classification (Sentiment Analysis) Named Entity Recognition (NER) Question Answering (QA) Text Summarization Machine Translation Text Generation (Dialogue) Fill-Mask / Language Modeling CV Computer Vision Image Classification Object Detection Image Segmentation Depth Estimation Image Generation Video Classification Keypoint Detection Audio & Multimodal Speech Recognition (ASR) Audio Classification Text-to-Speech (TTS) Image-Text Matching (VQA) Image Captioning Document Question Answering (Doc QA) Zero-Shot Classification

Core Principles of Transformer Architecture

Before using the library, understanding the underlying architecture will let you know why you tune parameters this way.

Overall Architecture: Encoder-Decoder

Three Major Model Families

Three Major Model Families: Architecture x Use Cases x Representative Models Encoder-Only Models Bidirectional attention, understands global context Suitable for: Classification, NER, QA (understanding tasks) Representatives: BERT, RoBERTa, ALBERT Chinese: BERT-wwm, MacBERT Feature: Can see all inputs simultaneously Input -> [CLS] represents global, [MASK] for pretraining Decoder-Only Models Causal attention (only looks left), autoregressive generation Suitable for: Text generation, dialogue, code generation Representatives: GPT series, LLaMA, Qwen Chinese: ChatGLM, Baichuan Feature: Predicts the next word one by one Input -> predict next token -> concatenate and predict again -> loop Encoder-Decoder model Sequence-to-sequence (Seq2Seq) task Suitable for: translation, summarization, question-answer generation Representatives: T5, BART, mT5 Chinese: mT5, PEGASUS-Chinese Features: encode input, decode to generate output Source sequence -> Encoder -> Decoder -> Target sequence

Installation and Environment Configuration

Installation

# 基础安装
pip install transformers

# 完整安装(包含训练依赖)
pip install transformers[torch]       # PyTorch 后端(推荐)
pip install transformers[tf-cpu]      # TensorFlow 后端
pip install transformers[flax]        # JAX/Flax 后端

# 常用配套库
pip install datasets          # HuggingFace 数据集库
pip install evaluate          # 模型评估指标
pip install accelerate        # 多GPU/混合精度训练
pip install peft              # 参数高效微调(LoRA等)
pip install tokenizers        # 高性能分词器
pip install sentencepiece     # 部分模型(T5/LLaMA)需要

# 验证安装
python -c "import transformers; print(transformers.__version__)"

Environment Variable Configuration

# 设置模型缓存目录(模型下载后缓存到此路径,默认 ~/.cache/huggingface)
export HF_HOME=/data/huggingface_cache

# 国内用户:使用镜像站加速下载(推荐 hf-mirror.com)
export HF_ENDPOINT=https://hf-mirror.com

# 离线模式(网络不可用时,只使用已缓存的模型)
export TRANSFORMERS_OFFLINE=1

# 禁用进度条(CI/CD 环境)
export DISABLE_TQDM=1

Example

# Can also be set in code
import os
os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"

# View current cache directory
from transformers.utils import TRANSFORMERS_CACHE
print(TRANSFORMERS_CACHE)

Pipeline: Run AI with Five Lines of Code

Pipeline is the highest-level abstraction in Transformers. It encapsulates model loading, preprocessing, inference, and post-processing entirely, allowing you to complete inference in just three to five lines of code.

Pipeline internal workflow Raw input Text / image Audio / other data Preprocessing Tokenizer tokenization Convert to token IDs Padding/Truncation Model inference Forward pass Output logits or hidden state vectors Post-processing Softmax/Argmax Decode tokens -> text Format output Result Label/score Generated text Pipeline automatically completes all steps; users only need to pass in raw input and get formatted results

Pipeline Quick Example Collection

Example

from transformers import pipeline

# 1. Sentiment analysis (text classification)
classifier = pipeline("sentiment-analysis")
result = classifier("I love using Hugging Face Transformers!")
# -> [{'label': 'POSITIVE', 'score': 0.9998}]

# 2. Text generation
generator = pipeline("text-generation", model="gpt2")
result = generator("Once upon a time in a land far away,",
    max_new_tokens=50, num_return_sequences=1, temperature=0.8)

# 3. Fill-in-the-blank (masked language model)
unmasker = pipeline("fill-mask", model="bert-base-uncased")
result = unmasker("The capital of France is [MASK].")
# -> [{'token_str': 'paris', 'score': 0.9823}, ...]

# 4. Named entity recognition (NER)
ner = pipeline("ner", aggregation_strategy="simple")
result = ner("My name is John and I work at Google in New York.")
# -> [{'entity_group': 'PER', 'word': 'John', 'score': 0.998}, ...]

# 5. Extractive question answering
qa = pipeline("question-answering")
result = qa(question="Who invented Python?",
    context="Python was created by Guido van Rossum in 1991.")
# -> {'answer': 'Guido van Rossum', 'score': 0.9887}

# 6. Text summarization
summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
result = summarizer(article, max_length=60, min_length=20)

# 7. Machine translation
translator = pipeline("translation", model="Helsinki-NLP/opus-mt-en-zh")
result = translator("Hello, how are you today?")
# -> [{'translation_text': 'Hello, how are you today?'}]

# 8. Zero-shot classification (no dedicated training needed)
zero_shot = pipeline("zero-shot-classification")
result = zero_shot("I love playing football",
    candidate_labels=["sports", "politics", "technology"])
# -> {'labels': ['sports', ...], 'scores': [0.972, ...]}

Advanced Pipeline Configuration

Example

import torch
from transformers import pipeline

# Specify GPU
pipe = pipeline("text-generation", model="gpt2", device=0)

# Specify precision (save memory)
pipe = pipeline("text-generation", model="meta-llama/Llama-2-7b-hf",
    torch_dtype=torch.float16, device_map="auto")

# Batch processing (increase throughput)
pipe = pipeline("sentiment-analysis", batch_size=32)
results = pipe(large_text_list)    # Automatic batch inference

# Large Text Chunking
asr = pipeline("automatic-speech-recognition",
    model="openai/whisper-large-v2",
    chunk_length_s=30, stride_length_s=5)
result = asr("long_audio.wav", return_timestamps=True)

In-Depth Analysis of Tokenizer

Tokenizer is the first step in NLP: converting raw text into a sequence of numbers that the model can understand.

Complete Tokenization Process

Complete Tokenization Process: Text -> Model Input Step 1 Raw Text: "Hello, I'm learning Transformers! It's great." Step 2 Tokenize: ["Hello", ",", "I", "'m", "learning", "Transform", "##ers", "!", ...] WordPiece/BPE subword tokenization: rare words are split (Transformers -> Transform + ##ers) Step 3 Add Special Tokens: ["[CLS]", "Hello", ",", "I", "'m", "learning", "Transform", "##ers", ... "[SEP]"] [CLS] classification token, [SEP] separator, [PAD] padding; different models have different special tokens Step 4 Convert to Token IDs: [101, 7592, 1010, 1045, 1005, 1049, 4083, 19081, 2121, ... 102] Each token maps to an integer index in the vocabulary and is fed into the model's Embedding layer.

Core Usage of Tokenizer

Example

from transformers import AutoTokenizer

# Load Tokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Encode in one step
encoding = tokenizer(
    "Hello, I'm learning Transformers!",
    return_tensors="pt",        # Return PyTorch tensor
    padding=True,               # Pad to the longest sequence
    truncation=True,            # Truncate when exceeding length
    max_length=128,             # Max length
)

print(encoding.keys())
# -> dict_keys(['input_ids', 'token_type_ids', 'attention_mask'])

print(encoding["input_ids"][0][:8])
# -> tensor([101, 7592, 1010, 1045, 1005, 1049, 4083, 19081])

print(encoding["attention_mask"][0][:8])
# -> tensor([1, 1, 1, 1, 1, 1, 1, 1]) # 1=real token, 0=padding

# Decode (ID -> Text)
decoded = tokenizer.decode(encoding["input_ids"][0], skip_special_tokens=True)
print(decoded)   # -> "hello, i'm learning transformers!"

# Batch encoding (automatic padding alignment)
texts = ["Short.", "This is a much longer sentence for testing."]
batch = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
print(batch["input_ids"].shape)   # -> torch.Size([2, 10])

# Vocabulary information
print(f"Vocabulary size: {tokenizer.vocab_size}")      # -> 30522
print(f"[CLS] ID: {tokenizer.cls_token_id}")     # -> 101
print(f"[SEP] ID: {tokenizer.sep_token_id}")     # -> 102
print(f"Max length: {tokenizer.model_max_length}")  # -> 512

Comparison of Common Tokenizer Types

Comparison of Three Mainstream Tokenization Algorithms Algorithm Principle Example Tokenization Representative Model BPE Byte Pair Encoding Merge high-frequency character pairs Learn optimal subword vocabulary "transformers" -> ["transform", "ers"] GPT / RoBERTa WordPiece Maximize language model probability ## Prefix-marked subwords "transformers" -> ["transform", "##ers"] BERT / DistilBERT SentencePiece Unigram / BPE variants Language-independent, direct processing Raw bytes, leading markers indicate word starts "transformers" -> ["_transform", "ers"] T5 / LLaMA / Qwen

Model Loading and Inference

AutoClass: Automatically Selecting the Correct Model Class

How AutoClass works: automatically matches the correct model architecture Model Name "bert-base-uncased" or local path AutoClass Load config.json Match architecture type AutoModel / AutoTokenizer / ... Automatically return the correct class BertForSequenceClassification GPT2LMHeadModel T5ForConditionalGeneration ... Common AutoClass quick reference: AutoTokenizer Auto tokenizer AutoModel Base model (outputs hidden states) AutoModelForSeqClass Text classification task AutoModelForCausalLM Text generation task

Example

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_name = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
    model_name, num_labels=2, torch_dtype=torch.float16, device_map="auto"
)

# Complete manual inference process
text = "Transformers is an amazing library!"

# 1. Encoding
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
inputs = {k: v.to(model.device) for k, v in inputs.items()}

# 2. Forward propagation
with torch.no_grad():
    outputs = model(**inputs)

# 3. Parse output
logits = outputs.logits                       # shape: [1, 2]
probs  = torch.softmax(logits, dim=-1)
pred   = torch.argmax(probs, dim=-1).item()

id2label = model.config.id2label              # {0: 'LABEL_0', 1: 'LABEL_1'}
print(f"Predicted class: {id2label[pred]}, confidence: {probs[0][pred]:.4f}")

Extracting Sentence Vectors

Example

from transformers import AutoModel, AutoTokenizer
import torch

model = AutoModel.from_pretrained("bert-base-uncased")
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

def get_sentence_embedding(text: str) -> torch.Tensor:
    inputs = tokenizer(text, return_tensors="pt", max_length=512, truncation=True)
    with torch.no_grad():
        outputs = model(**inputs)
    # Apply mean pooling to all tokens
    token_embeddings = outputs.last_hidden_state          # [1, seq_len, 768]
    attention_mask = inputs["attention_mask"].unsqueeze(-1)
    mean_embedding = (token_embeddings * attention_mask).sum(1) / attention_mask.sum(1)
    return mean_embedding  # [1, 768]

vec = get_sentence_embedding("Hello world")
print(vec.shape)  # -> torch.Size([1, 768])

Practical Guide to Ten Common Tasks

Text Classification (Sentiment Analysis)

Example

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_name = "cardiffnlp/twitter-roberta-base-sentiment-latest"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

def predict_sentiment(texts):
    inputs = tokenizer(texts, return_tensors="pt", padding=True,
                       truncation=True, max_length=512)
    with torch.no_grad():
        logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)
    results = []
    for i, t in enumerate(texts):
        pid = probs[i].argmax().item()
        results.append({"text": t, "label": model.config.id2label[pid],
                         "score": round(probs[i][pid].item(), 4)})
    return results

print(predict_sentiment(["I love this!", "This is terrible."]))
# -> [{'text': 'I love this!', 'label': 'positive', 'score': 0.9756}, ...]

Text Generation (Dialogue / Continuation)

Example

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_name = "Qwen/Qwen2-1.5B-Instruct"     # Alibaba Tongyi Qianwen (supports Chinese)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype=torch.float16, device_map="auto"
)

messages = [
    {"role": "system", "content": "You are a helpful AI assistant."},
    {"role": "user", "content": "Please explain what a Transformer is in three sentences."},
]
text = tokenizer.apply_chat_template(messages, tokenize=False,
                                      add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

with torch.no_grad():
    output_ids = model.generate(
        **inputs, max_new_tokens=300, temperature=0.7, top_p=0.9,
        do_sample=True, repetition_penalty=1.1,
        pad_token_id=tokenizer.eos_token_id,
    )

new_tokens = output_ids[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(new_tokens, skip_special_tokens=True)
print(response)

Named Entity Recognition (NER)

Example

from transformers import pipeline

ner = pipeline("ner", model="dslim/bert-base-NER", aggregation_strategy="simple")
result = ner("Elon Musk founded SpaceX in 2002 and Tesla Motors in 2003.")
for entity in result:
    print(f"{entity['word']:<20} -> {entity['entity_group']} ({entity['score']:.3f})")

# Chinese NER
ner_cn = pipeline("ner", model="hfl/chinese-bert-wwm-ext-ner-msra",
                   aggregation_strategy="simple")
result = ner_cn("Xiao Ming studied at Peking University and later went to work at Alibaba.")

Machine Translation

Example

from transformers import pipeline

# English to Chinese
translator = pipeline("translation", model="Helsinki-NLP/opus-mt-en-zh")
result = translator("Artificial intelligence is transforming the world.")
print(result[0]["translation_text"])   # -> Artificial intelligence is changing the world.

# Chinese to English
translator_zh = pipeline("translation", model="Helsinki-NLP/opus-mt-zh-en")
result = translator_zh("Artificial intelligence is changing the world.")
print(result[0]["translation_text"])

Text Summarization

Example

from transformers import pipeline

summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
result = summarizer(long_text, max_length=80, min_length=30,
                     do_sample=False, no_repeat_ngram_size=3)
print(result[0]["summary_text"])

Fine-tuning

Fine-tuning adapts a pretrained model to your specific task and data, and is the most important application scenario of Transformers.

Overview of the Fine-tuning Process

Complete fine-tuning process with Transformers Prepare data Dataset loading HF datasets or custom CSV Tokenize encoding Padding alignment DataLoader Key: data quality > Data quantity Load model AutoModelForTask Specify num_labels Freeze bottom-layer parameters (optional) Only fine-tune the top layers Tip: small dataset Freeze the bottom layers first Training configuration TrainingArguments Learning rate: 2e-5 Batch:32 Epochs:3~5 Warmup steps Weight decay Tip: Learning rate Most important Training & Evaluation Trainer.train() Monitor loss curve Validation set evaluation Early stopping Save best ckpt Note: Overfitting Compare baseline Save save_model() save_pretrained Push to Hub Production deployment Quantized inference ONNX export Using HuggingFace Trainer can greatly simplify the training pipeline, automatically handling gradient accumulation, mixed precision, distributed training, etc.

Complete Fine-tuning Example: Text Classification

Example

from datasets import load_dataset
from transformers import (
    AutoTokenizer, AutoModelForSequenceClassification,
    TrainingArguments, Trainer, DataCollatorWithPadding,
    EarlyStoppingCallback,
)
import evaluate, numpy as np

# 1. Load dataset
dataset = load_dataset("imdb")   # HF Hub public dataset

# 2. Tokenizer + preprocessing
MODEL_NAME = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)

def tokenize_fn(examples):
    return tokenizer(examples["text"], truncation=True, max_length=512)

tokenized_ds = dataset.map(tokenize_fn, batched=True,
                           remove_columns=["text"])
tokenized_ds = tokenized_ds.rename_column("label", "labels")
tokenized_ds.set_format("torch")
data_collator = DataCollatorWithPadding(tokenizer=tokenizer)

# 3. Load model
model = AutoModelForSequenceClassification.from_pretrained(
    MODEL_NAME, num_labels=2,
    id2label={0: "NEGATIVE", 1: "POSITIVE"},
    label2id={"NEGATIVE": 0, "POSITIVE": 1},
)

# 4. Evaluation metrics
accuracy = evaluate.load("accuracy")
f1 = evaluate.load("f1")
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    preds = np.argmax(logits, axis=-1)
    return {"accuracy": accuracy.compute(predictions=preds, references=labels)["accuracy"],
            "f1": f1.compute(predictions=preds, references=labels, average="binary")["f1"]}

# 5. Training arguments
training_args = TrainingArguments(
    output_dir="./results", num_train_epochs=3,
    per_device_train_batch_size=16, per_device_eval_batch_size=32,
    gradient_accumulation_steps=2, learning_rate=2e-5,
    weight_decay=0.01, warmup_ratio=0.1,
    evaluation_strategy="steps", eval_steps=500,
    save_strategy="steps", save_steps=500,
    load_best_model_at_end=True, metric_for_best_model="f1",
    fp16=True, logging_steps=100, seed=42,
)

# 6. Create Trainer and train
trainer = Trainer(
    model=model, args=training_args,
    train_dataset=tokenized_ds["train"],
    eval_dataset=tokenized_ds["test"],
    tokenizer=tokenizer, data_collator=data_collator,
    compute_metrics=compute_metrics,
    callbacks=[EarlyStoppingCallback(early_stopping_patience=3)],
)
trainer.train()

# 7. Evaluate and save
eval_result = trainer.evaluate()
print(f"Accuracy: {eval_result['eval_accuracy']:.4f}")
print(f"F1: {eval_result['eval_f1']:.4f}")

trainer.save_model("./my-sentiment-model")
tokenizer.save_pretrained("./my-sentiment-model")

LoRA Parameter-Efficient Fine-tuning (Recommended)

LoRA principle: only train low-rank matrices, freeze the original weights Full Fine-tuning W + ΔW Update all weight matrices Parameter count: all (e.g., train all 7B parameters) GPU memory requirement: extremely high (requires 4x model size) Cost: expensive, requires many GPUs LoRA fine-tuning W (frozen) Not trained + B x A r << d (low-rank matrix) Only train here Parameter count: 0.1% to 1% of the original GPU memory requirement: low (only stores low-rank matrix gradients) Cost: a single 24GB GPU can fine-tune a 7B model LoRA decomposes ΔW into two low-rank matrices, where r is much smaller than d (typically r=8~64), greatly reducing training cost.

Example

from peft import LoraConfig, get_peft_model, TaskType
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments
import torch

model_name = "meta-llama/Llama-2-7b-hf"
model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype=torch.float16, device_map="auto",
    load_in_4bit=True,        # 4-bit quantized loading (QLoRA), further saves GPU memory
)

# Configure LoRA
lora_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM, r=16, lora_alpha=32,
    lora_dropout=0.05,
    target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# -> trainable: 4,194,304 || all: 6,742,609,920 || 0.0622%

LoRA advantages: only trains less than 1% of parameters, reduces GPU memory by 60-70%, is 2-3x faster, weight files are only a few MB, and multiple LoRA adapters can be saved for the same base model for different tasks.


Model Saving, Loading, and Publishing

Example

# Save locally
model.save_pretrained("./my-model")
tokenizer.save_pretrained("./my-model")

# Load locally
model = AutoModelForSequenceClassification.from_pretrained("./my-model")
tokenizer = AutoTokenizer.from_pretrained("./my-model")

# Publish to HuggingFace Hub
from huggingface_hub import login
login(token="your_hf_token")   # huggingface.co/settings/tokens
model.push_to_hub("your-username/my-sentiment-model")
tokenizer.push_to_hub("your-username/my-sentiment-model")

# Publish directly via Trainer
training_args = TrainingArguments(
    output_dir="your-username/my-model",
    push_to_hub=True, hub_strategy="every_save",
)

Performance Optimization Tips

Overview of Inference Acceleration

Inference Acceleration Tech Stack: From Simple to Deep Optimization Level 1 - Zero-Cost Optimization torch.no_grad() inference | fp16/bf16 half precision | batch inference | device_map="auto" automatic device assignment Speedup: 1.5~2x Level 2 - Quantization (Model Compression) bitsandbytes 4-bit/8-bit quantization | GPTQ (post-training quantization) | AWQ (activation-aware quantization) Speedup: 2~4x, VRAM halved Level 3 - Compilation and Runtime Optimization torch.compile()(PyTorch 2.0)| FlashAttention-2 | xFormers | Optimum(TensorRT/ONNX) Speedup: 3~10x Level 4 - Specialized Inference Engines vLLM (LLM high-throughput inference) | TGI | TensorRT-LLM | llama.cpp (CPU) Production-grade, highest performance

Example

# 4-bit quantized loading (13B model only needs ~7GB VRAM)
from transformers import BitsAndBytesConfig, AutoModelForCausalLM
import torch

quant_config = BitsAndBytesConfig(
    load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4",
)
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-2-13b-hf",
    quantization_config=quant_config, device_map="auto",
)

# FlashAttention-2 acceleration (requires pip install flash-attn)
model = AutoModelForCausalLM.from_pretrained(
    "mistralai/Mistral-7B-v0.1",
    attn_implementation="flash_attention_2",
    torch_dtype=torch.bfloat16, device_map="auto",
)

# torch.compile (PyTorch 2.0+)
model = torch.compile(model, mode="reduce-overhead")

Troubleshooting Common Issues

Quick Reference for Common Errors and Solutions Error Message Cause and Solution CUDA out of memory Out of VRAM Reduce batch_size; add gradient_accumulation_steps; use fp16/4-bit; use a smaller model OSError: model not found Model name error or network issue Check spelling; set HF_ENDPOINT mirror; use local path if already downloaded ValueError: num_labels mismatch Explicitly specify num_labels=your number of classes when loading the model tensors on different devices inputs = {k: v.to(model.device) for k, v in inputs.items()} loss = NaN / loss not decreasing Check learning rate (too large causes NaN); check labels range (0~N-1); add gradient clipping slow tokenizer / slow speed pip install tokenizers to install the fast Rust version; use use_fast=True (default)

Summary and Learning Path

Transformers Learning Path and Core Knowledge Points Phase 1 Getting Started (1 week) pip install and Environment Setup Run 5 Tasks with Pipeline Understand the Tokenizer Workflow Manual Inference with AutoModel Understand Encoder/Decoder Understand Model Output Structure Phase 2 Advanced (2 weeks) Load Models from the Hub Hands-on Fine-tuning for Text Classification Master the Trainer API Custom Dataset Processing Evaluation with compute_metrics Save and Publish to the Hub Phase 3 Advanced (3 weeks) LoRA / PEFT Fine-tuning Quantization (4-bit/8-bit) Using Multimodal Models Custom Training Loop Accelerate Multi-GPU FlashAttention Acceleration Phase 4 Expert (Ongoing) Custom Model Architecture Pretraining from Scratch vLLM Production Deployment RLHF/DPO Alignment Contributing Open-Source Models Research Cutting-Edge Papers Each phase should include hands-on practice: find a real dataset and run through the complete workflow of training -> evaluation -> publishing

Key API Quick Reference

Example

# 1. Load Tokenizer
tokenizer = AutoTokenizer.from_pretrained("model_name")

# 2. Encode Text
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)

# 3. Load Model
model = AutoModelForSequenceClassification.from_pretrained("model_name", num_labels=N)

# 4. Inference
with torch.no_grad():
    outputs = model(**inputs)

# 5. Pipeline
pipe = pipeline("task_name", model="model_name")

# 6. Training Configuration
args = TrainingArguments(output_dir="./out", num_train_epochs=3, learning_rate=2e-5)

# 7. Training
trainer = Trainer(model=model, args=args, train_dataset=ds, compute_metrics=fn)

# 8. Save
model.save_pretrained("./my-model")
tokenizer.save_pretrained("./my-model")

# 9. Dataset
dataset = load_dataset("dataset_name")
dataset = dataset.map(tokenize_fn, batched=True)

# 10. LoRA
config = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj","v_proj"])
model = get_peft_model(model, config)

Reference Resources

Other Extensions