Computer Vision AI
Of the information humans acquire, about 80% comes from vision.
When you see a photo, you can immediately recognize how many people are in it, what they are doing, and whether the background is indoors or outdoors.
But for a computer, this photo is just a pile of numbers—each pixel is represented by three values for red, green, and blue, and that's all.
Computer Vision (CV) is the technology that enables computers to understand images.
From facial unlock on phones to road condition recognition in autonomous driving, and to lesion analysis in medical imaging, computer vision has become ubiquitous.
This module will take you from basic convolutional neural networks all the way to the latest vision-language models.
Learning path: Convolutional Neural Networks → Vision Transformer → Object Detection → Image Segmentation → Diffusion Models → CLIP → Vision-Language Models. Each step comes with runnable code examples.
Convolutional Neural Networks (CNN)
CNN is a foundational technology in computer vision. It mimics the way the human visual cortex works, extracting image features through local receptive fields.
Principles of Convolution Operation
The core idea of convolution is:Use a small sliding window (convolution kernel) to scan the image and extract local features.。
For example, a 3×3 kernel looks at 9 pixels at a time, computes their weighted sum, and produces an output value.
This process is repeated; the kernel slides over the image from left to right and top to bottom, ultimately generating a "feature map".
Example
# Implement the simplest convolution operation using NumPy
# Demonstrate how convolution extracts edge features
# ============================================
import numpy as np
def simple_convolution(image: np.ndarray, kernel: np.ndarray) -> np.ndarray:
"""
Implement the most basic 2D convolution operation (without padding and stride)
Parameters:
image: input image (H, W), single-channel grayscale
kernel: convolution kernel (kH, kW)
Returns:
Feature map after convolution
"""
# Get the dimensions of the image and the kernel
img_h, img_w = image.shape
kernel_h, kernel_w = kernel.shape
# Compute the size of the output feature map
# Output size = input size - kernel size + 1
out_h = img_h - kernel_h + 1
out_w = img_w - kernel_w + 1
# Initialize the output feature map
output = np.zeros((out_h, out_w))
# Slide the convolution kernel to compute
for i in range(out_h):
for j in range(out_w):
# Extract the local region of the image corresponding to the kernel
region = image[i:i+kernel_h, j:j+kernel_w]
# Multiply corresponding elements and sum (this is the convolution operation)
output[i, j] = np.sum(region * kernel)
return output
# Create a simple test image: a white square in the middle, black around it
# Shape: 8×8 grayscale image
test_image = np.array([
[0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 1, 1, 1, 1, 0, 0],
[0, 0, 1, 1, 1, 1, 0, 0],
[0, 0, 1, 1, 1, 1, 0, 0],
[0, 0, 1, 1, 1, 1, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0],
])
print("Original image:")
print(test_image)
# Define an edge detection kernel (simplified Sobel operator)
# This kernel can detect vertical edges
edge_kernel = np.array([
[-1, 0, 1],
[-2, 0, 2],
[-1, 0, 1],
])
# Perform convolution
feature_map = simple_convolution(test_image, edge_kernel)
print("\nFeature map after convolution (vertical edges detected):")
print(np.round(feature_map, 2))
The key to the convolution operation is:The parameters of the convolution kernel are learned, not manually designed.。
During training, the model automatically adjusts the kernel values, allowing it to extract features useful for the task.
Pooling
Pooling compresses the size of the feature map, reduces computation, while preserving important features.
The most commonly used is max pooling: divide the feature map into several small blocks, and each block retains only the maximum value.
Example
# Implement the max pooling operation
# ============================================
def max_pooling(feature_map: np.ndarray, pool_size: int = 2) -> np.ndarray:
"""
Implement the max pooling operation
Parameters:
feature_map: input feature map (H, W)
pool_size: pooling window size, default 2×2
Returns:
The pooled feature map
"""
h, w = feature_map.shape
# Calculate output size
out_h = h // pool_size
out_w = w // pool_size
output = np.zeros((out_h, out_w))
for i in range(out_h):
for j in range(out_w):
# Extract the pooling window region
region = feature_map[
i*pool_size:(i+1)*pool_size,
j*pool_size:(j+1)*pool_size
]
# Take the maximum value
output[i, j] = np.max(region)
return output
# Perform pooling using the feature map from earlier
print("Feature map before pooling:")
print(feature_map)
pooled = max_pooling(feature_map, pool_size=2)
print("\n"Feature map after 2×2 max pooling:")
print(pooled)
Classic Architecture Comparison
In the history of CNN development, there are several milestone architectures, each representing design ideas from different periods.
| Architecture | Year | Core innovation | Characteristics |
|---|---|---|---|
| AlexNet | 2012 | ReLU activation, Dropout, data augmentation | The pioneer of the deep learning era, first to significantly outperform traditional methods on ImageNet. |
| VGG | 2014 | Uniform use of 3×3 small convolution kernels | Simple and elegant structure, easy to understand and implement |
| ResNet | 2015 | Residual connection (Skip Connection) | Solves the vanishing gradient problem in deep networks, allowing networks to have hundreds of layers. |
| EfficientNet | 2019 | Compound scaling method | Scales depth, width, and resolution simultaneously, with extremely high parameter efficiency. |
Residual Connection (Skip Connection)
The core innovation of ResNet is the residual connection, which solves the problem that "the deeper the network, the harder it is to train."
The learning goal of a traditional network layer is to directly learn the mapping from input x to output y.
The learning goal of a residual network is:Learn the difference (residual) between output y and input x.。
Formula: y = F(x) + x, where F(x) is the residual to be learned.
Intuition of residual connection: letting the network learn "how much to change based on the existing state" is much easier than letting it "learn the complete mapping from scratch."
Example
# Implement a complete CNN image classifier with PyTorch
# Includes convolution, pooling, residual connections
# ============================================
import torch
import torch.nn as nn
import torch.nn.functional as F
class ResidualBlock(nn.Module):
"""A simple residual block"""
def __init__(self, in_channels: int, out_channels: int, stride: int = 1):
super().__init__()
# First convolutional layer
self.conv1 = nn.Conv2d(
in_channels, out_channels,
kernel_size=3, stride=stride, padding=1, bias=False
)
self.bn1 = nn.BatchNorm2d(out_channels)
# Second convolutional layer
self.conv2 = nn.Conv2d(
out_channels, out_channels,
kernel_size=3, stride=1, padding=1, bias=False
)
self.bn2 = nn.BatchNorm2d(out_channels)
# Shortcut connection (for matching dimensional changes)
self.shortcut = nn.Sequential()
if stride != 1 or in_channels != out_channels:
self.shortcut = nn.Sequential(
nn.Conv2d(
in_channels, out_channels,
kernel_size=1, stride=stride, bias=False
),
nn.BatchNorm2d(out_channels)
)
def forward(self, x: torch.Tensor) -> torch.Tensor:
# Main path: conv → BN → ReLU → conv → BN
out = F.relu(self.bn1(self.conv1(x)))
out = self.bn2(self.conv2(out))
# Residual connection: add the input (or the input after dimension transformation)
out += self.shortcut(x)
# Final ReLU
out = F.relu(out)
return out
class SimpleCNN(nn.Module):
"""Simple CNN for image classification (with residual connections)"""
def __init__(self, num_classes: int = 10):
super().__init__()
# Initial convolutional layer
self.conv1 = nn.Conv2d(3, 32, kernel_size=3, stride=1, padding=1)
self.bn1 = nn.BatchNorm2d(32)
# Residual layer
self.layer1 = ResidualBlock(32, 32)
self.layer2 = ResidualBlock(32, 64, stride=2)
self.layer3 = ResidualBlock(64, 64)
# Global average pooling
self.global_pool = nn.AdaptiveAvgPool2d((1, 1))
# Classification head
self.fc = nn.Linear(64, num_classes)
def forward(self, x: torch.Tensor) -> torch.Tensor:
# x shape: (batch_size, 3, 32, 32)
x = F.relu(self.bn1(self.conv1(x)))
x = self.layer1(x) # (batch_size, 32, 32, 32)
x = self.layer2(x) # (batch_size, 64, 16, 16)
x = self.layer3(x) # (batch_size, 64, 16, 16)
x = self.global_pool(x) # (batch_size, 64, 1, 1)
x = x.view(x.size(0), -1) # (batch_size, 64)
x = self.fc(x) # (batch_size, num_classes)
return x
# Create the model and test
model = SimpleCNN(num_classes=10)
print("Model structure:")
print(model)
# Test forward pass
# Create a fake batch: 4 images, each 3×32×32 (RGB 32×32)
test_input = torch.randn(4, 3, 32, 32)
test_output = model(test_input)
print(f"\nInput shape: {test_input.shape}")
print(f"Output shape: {test_output.shape} (batch_size, num_classes)")
Vision Transformer(ViT)
In 2020, Google published the paper "An Image is Worth 16x16 Words", bringing the Transformer architecture from the language domain to the vision domain.
Image Patching (Patch Embedding)
Transformers process sequence data, but images are 2D grid data.
ViT's solution is simple:Cut the image into small patches, and treat each patch as a word。
For example, a 224×224 image cut into 16×16 patches yields 14×14 = 196 patches.
Each patch is flattened into a 1D vector and projected into an embedding via a linear layer, which can then be fed into the Transformer.
Example
# Implement Patch Embedding (image patching)
# ============================================
class PatchEmbedding(nn.Module):
"""
Segment the image into patches and embed them
Parameters:
img_size: input image size (square)
patch_size: size of each patch (square)
in_channels: number of input channels (3 for RGB)
embed_dim: output embedding dimension
"""
def __init__(
self,
img_size: int = 224,
patch_size: int = 16,
in_channels: int = 3,
embed_dim: int = 768
):
super().__init__()
self.img_size = img_size
self.patch_size = patch_size
# Calculate the number of patches: (224/16)^2 = 14^2 = 196
self.num_patches = (img_size // patch_size) ** 2
# Use convolution to implement patch partitioning and embedding
# This is equivalent to: cut patches → flatten → linear projection
self.proj = nn.Conv2d(
in_channels, embed_dim,
kernel_size=patch_size, stride=patch_size
)
def forward(self, x: torch.Tensor) -> torch.Tensor:
# x shape: (batch_size, in_channels, img_size, img_size)
x = self.proj(x) # (batch_size, embed_dim, num_patches^0.5, num_patches^0.5)
x = x.flatten(2) # (batch_size, embed_dim, num_patches)
x = x.transpose(1, 2) # (batch_size, num_patches, embed_dim)
return x
# Test Patch Embedding
patch_embed = PatchEmbedding(img_size=224, patch_size=16, embed_dim=768)
test_image = torch.randn(1, 3, 224, 224) # 1 RGB image of 224×224
patches = patch_embed(test_image)
print(f"Input image shape: {test_image.shape}")
print(f"Patch sequence shape: {patches.shape} (batch_size, num_patches, embed_dim)")
ViT Full Architecture
The complete ViT architecture also requires adding a class token and positional encoding.
The class token is a learnable vector that serves to "aggregate" information from all patches.
Positional encoding is also learnable; it tells the model "where each patch is located in the image".
Example
# Implement a simplified Vision Transformer
# ============================================
class MultiHeadAttention(nn.Module):
"""Multi-head self-attention"""
def __init__(self, dim: int, num_heads: int):
super().__init__()
self.num_heads = num_heads
self.head_dim = dim // num_heads
assert self.head_dim * num_heads == dim, "Dimension must be divisible by the number of heads"
# Q, K, V projection matrices
self.qkv = nn.Linear(dim, dim * 3)
# Output projection
self.proj = nn.Linear(dim, dim)
def forward(self, x: torch.Tensor) -> torch.Tensor:
batch_size, seq_len, dim = x.shape
# Compute Q, K, V
qkv = self.qkv(x) # (batch_size, seq_len, dim*3)
qkv = qkv.reshape(
batch_size, seq_len, 3, self.num_heads, self.head_dim
) # (batch_size, seq_len, 3, num_heads, head_dim)
q, k, v = qkv.unbind(2) # Each shape: (batch_size, seq_len, num_heads, head_dim)
# Rearrange dimensions to compute attention
q = q.transpose(1, 2) # (batch_size, num_heads, seq_len, head_dim)
k = k.transpose(1, 2)
v = v.transpose(1, 2)
# Compute attention scores
attention = q @ k.transpose(-2, -1) / (self.head_dim ** 0.5)
attention = attention.softmax(dim=-1)
# Aggregate values
out = attention @ v # (batch_size, num_heads, seq_len, head_dim)
out = out.transpose(1, 2) # (batch_size, seq_len, num_heads, head_dim)
out = out.flatten(2) # (batch_size, seq_len, dim)
# Output projection
out = self.proj(out)
return out
class TransformerBlock(nn.Module):
"""A Transformer block"""
def __init__(self, dim: int, num_heads: int, mlp_ratio: float = 4.0):
super().__init__()
self.norm1 = nn.LayerNorm(dim)
self.attn = MultiHeadAttention(dim, num_heads)
self.norm2 = nn.LayerNorm(dim)
# MLP
mlp_hidden = int(dim * mlp_ratio)
self.mlp = nn.Sequential(
nn.Linear(dim, mlp_hidden),
nn.GELU(),
nn.Linear(mlp_hidden, dim)
)
def forward(self, x: torch.Tensor) -> torch.Tensor:
# Residual connection 1
x = x + self.attn(self.norm1(x))
# Residual connection 2
x = x + self.mlp(self.norm2(x))
return x
class SimpleViT(nn.Module):
"""Simplified Vision Transformer"""
def __init__(
self,
img_size: int = 224,
patch_size: int = 16,
in_channels: int = 3,
num_classes: int = 10,
embed_dim: int = 192,
depth: int = 6,
num_heads: int = 6
):
super().__init__()
# Patch embedding
self.patch_embed = PatchEmbedding(
img_size, patch_size, in_channels, embed_dim
)
num_patches = self.patch_embed.num_patches
# Class token
self.cls_token = nn.Parameter(torch.zeros(1, 1, embed_dim))
# Positional encoding
self.pos_embed = nn.Parameter(
torch.zeros(1, num_patches + 1, embed_dim)
)
# Transformer layers
self.blocks = nn.ModuleList([
TransformerBlock(embed_dim, num_heads)
for _ in range(depth)
])
# Final LayerNorm and classification head
self.norm = nn.LayerNorm(embed_dim)
self.head = nn.Linear(embed_dim, num_classes)
def forward(self, x: torch.Tensor) -> torch.Tensor:
batch_size = x.shape[0]
# Patch embedding
x = self.patch_embed(x) # (batch_size, num_patches, embed_dim)
# Add class token
cls_tokens = self.cls_token.expand(batch_size, -1, -1)
x = torch.cat((cls_tokens, x), dim=1) # (batch_size, 1+num_patches, embed_dim)
# Add positional encoding
x = x + self.pos_embed
# Pass through Transformer layers
for block in self.blocks:
x = block(x)
# Use only the class token for classification
x = self.norm(x)
cls_output = x[:, 0] # Take the first position (class token)
x = self.head(cls_output)
return x
# Create ViT model and test
vit = SimpleViT(
img_size=32, # Use smaller images for easier testing
patch_size=8, # 8×8 patch
embed_dim=128,
depth=4,
num_heads=4,
num_classes=10
)
print("Simplified ViT architecture:")
print(vit)
test_input = torch.randn(2, 3, 32, 32) # 2 32×32 images
test_output = vit(test_input)
print(f"\nInput shape: {test_input.shape}")
print(f"Output shape: {test_output.shape}")
CNN vs ViT Comparison
Two architectures each have their own pros and cons; which one you choose depends on your specific scenario.
| Characteristics | CNN | ViT |
|---|---|---|
| Inductive bias | Strong (locality, translation invariance) | Weak (learns from data) |
| Data requirements | Works with small data | Needs large amounts of data to show its advantages |
| Computational efficiency | Usually faster, mature hardware optimization | Large parameter count, computationally intensive |
| Interpretability | Easy to visualize convolution kernels | Attention maps are interpretable, but the overall model is harder to interpret |
| Long-range dependencies | Require multiple layers to establish | Naturally supports global receptive field |
| Applicable scenarios | Limited data, high real-time requirements | Large datasets, complex visual tasks |
Rule of thumb: use CNN (ResNet, EfficientNet) when data is limited; consider ViT when data reaches millions or more. There are also hybrid architectures (such as ConvNeXt) that combine the advantages of both.
Object Detection
Image classification answers "what is in this image," while object detection answers "what is in this image and where are they located."
The output of object detection is a series of bounding boxes, each corresponding to an object category.
Two-Stage Detection: R-CNN Series
The idea of two-stage detection is:First find regions that may contain objects, then classify these regions。
First stage: generate region proposals, possibly hundreds of boxes indicating "something might be here."
Second stage: extract features from each candidate region, determine what object it is, and adjust the box position.
Evolution of the R-CNN family:
| Model | Key improvements | Speed |
|---|---|---|
| R-CNN | Replaces traditional features with CNN | Slow (tens of seconds per image) |
| Fast R-CNN | ROI Pooling, shared feature computation | Medium (a few seconds per image) |
| Faster R-CNN | RPN(Region Proposal Network) | Fast (hundreds of milliseconds per image) |
| Mask R-CNN | Adds a segmentation branch | Medium |
One-Stage Detection: YOLO Series
The idea of one-stage detection is:Directly perform dense prediction on the image without generating candidate regions separately。
YOLO (You Only Look Once) is a representative of one-stage detection, extremely fast, and suitable for real-time applications.
YOLO divides the image into an S×S grid, with each grid cell responsible for predicting objects whose center point falls within that cell.
Each grid cell predicts: B bounding boxes (with confidence) and probabilities for C classes.
End-to-End Detection: DETR
DETR (DEtection TRansformer) is a new paradigm proposed by Facebook AI, using Transformer to turn object detection into a direct sequence prediction problem.
Its biggest feature is: no need for non-maximum suppression (NMS), no need for anchor boxes (Anchor), the output is the final result.
Hands-On: YOLOv11 Object Detection
Now we use Ultralytics YOLOv11 to create a complete object detection example.
Example
# YOLOv11 Object Detection Hands-on
# Includes model loading, detection, visualization
# ============================================
import torch
import numpy as np
from PIL import Image, ImageDraw, ImageFont
import matplotlib.pyplot as plt
class SimpleYOLO:
"""Simplified YOLO demo (using pretrained model)"""
def __init__(self, model_name: str = "yolo11n.pt"):
"""
Initialize YOLO model
Parameters:
model_name: model name, yolo11n.pt (nano, small), yolo11s.pt (small), etc.
"""
self.model_name = model_name
self.model = None
# 80 categories of the COCO dataset
self.coco_names = [
'person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus',
'train', 'truck', 'boat', 'traffic light', 'fire hydrant',
'stop sign', 'parking meter', 'bench', 'bird', 'cat', 'dog',
'horse', 'sheep', 'cow', 'elephant', 'bear', 'zebra', 'giraffe',
'backpack', 'umbrella', 'handbag', 'tie', 'suitcase', 'frisbee',
'skis', 'snowboard', 'sports ball', 'kite', 'baseball bat',
'baseball glove', 'skateboard', 'surfboard', 'tennis racket',
'bottle', 'wine glass', 'cup', 'fork', 'knife', 'spoon', 'bowl',
'banana', 'apple', 'sandwich', 'orange', 'broccoli', 'carrot',
'hot dog', 'pizza', 'donut', 'cake', 'chair', 'couch',
'potted plant', 'bed', 'dining table', 'toilet', 'tv', 'laptop',
'mouse', 'remote', 'keyboard', 'cell phone', 'microwave', 'oven',
'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
'scissors', 'teddy bear', 'hair drier', 'toothbrush'
]
# Color for each category
self.colors = np.random.randint(0, 255, (80, 3))
def load_model(self):
"""Load pretrained model (here using a simulated implementation to demonstrate the concept)"""
print(f"Loading model {self.model_name}...")
print("Note: In actual use, please install ultralytics and call:")
print(" from ultralytics import YOLO")
print(" model = YOLO('yolo11n.pt')")
self.model = "pretrained_model_loaded"
print("Model loaded successfully!")
return self
def predict_dummy(self, image: np.ndarray) -> list:
"""
Simulated detection results (for demonstration)
In actual use, call model(image)
Returns:
Detection result list, each element is a dictionary:
{'bbox': [x1, y1, x2, y2], 'class_id': int, 'score': float, 'class_name': str}
"""
# Here we simulate detecting several objects
h, w = image.shape[:2]
dummy_results = [
{'bbox': [w*0.1, h*0.2, w*0.3, h*0.6], 'class_id': 0, 'score': 0.92, 'class_name': 'person'},
{'bbox': [w*0.35, h*0.4, w*0.7, h*0.8], 'class_id': 2, 'score': 0.85, 'class_name': 'car'},
{'bbox': [w*0.75, h*0.3, w*0.9, h*0.55], 'class_id': 16, 'score': 0.78, 'class_name': 'dog'},
]
return dummy_results
def visualize(self, image: np.ndarray, results: list) -> np.ndarray:
"""
Draw detection boxes and labels on the image
Parameters:
image: original image (H, W, 3)
results: detection result list
Returns:
Annotated image
"""
if isinstance(image, np.ndarray):
image = Image.fromarray(image)
draw = ImageDraw.Draw(image)
for result in results:
bbox = result['bbox']
class_id = result['class_id']
score = result['score']
class_name = result['class_name']
color = tuple(self.colors[class_id % 80].tolist())
# Draw box
draw.rectangle(bbox, outline=color, width=3)
# Label background
label = f"{class_name} {score:.2f}"
try:
font = ImageFont.truetype("arial.ttf", 16)
except:
font = ImageFont.load_default()
# Calculate text size
bbox_label = draw.textbbox((0, 0), label, font=font)
text_w = bbox_label[2] - bbox_label[0]
text_h = bbox_label[3] - bbox_label[1]
# Draw label background
label_bg = [bbox[0], bbox[1]-text_h-5, bbox[0]+text_w+5, bbox[1]]
draw.rectangle(label_bg, fill=color)
# Draw text
draw.text((bbox[0]+2, bbox[1]-text_h-3), label, fill=(255,255,255), font=font)
return np.array(image)
# Create simulated image (for demonstration)
def create_test_image() -> np.ndarray:
"""Create a test image"""
h, w = 480, 640
image = np.zeros((h, w, 3), dtype=np.uint8)
# Gradient background
image[:, :, 0] = np.linspace(100, 200, w)[np.newaxis, :]
image[:, :, 1] = np.linspace(150, 100, h)[:, np.newaxis]
image[:, :, 2] = 180
return image
# Demonstrate YOLO detection process
print("=" * 60)
print("EXAMPLE YOLOv11 Object Detection Demo")
print("=" * 60)
# Initialize
detector = SimpleYOLO().load_model()
# Create test image
test_image = create_test_image()
print(f"\nInput image size: {test_image.shape}")
# Simulated detection
results = detector.predict_dummy(test_image)
print(f"\n"Detected {len(results)} objects:")
for i, r in enumerate(results, 1):
print(f" {i}. {r['class_name']} (Confidence: {r['score']:.2f})")
# Visualization
visualized = detector.visualize(test_image, results)
print(f"\nVisualization complete, output image size: {visualized.shape}")
# ============================================
# Actual code for using YOLOv11 (requires installing ultralytics)
# ============================================
print("\n" + "=" * 60)
print("Example code for actually using YOLOv11:")
print("=" * 60)
print("""
# Install: pip install ultralytics
from ultralytics import YOLO
# 1. Load model
model = YOLO('yolo11n.pt') # You can also use 'yolo11s.pt', 'yolo11m.pt', etc.
# 2. Predict
results = model('test_image.jpg') # Can be an image path, video, or 0 (webcam)
# 3. Process results
for result in results:
# Detection boxes
boxes = result.boxes
# Segmentation masks (if using a segmentation model)
masks = result.masks
# Display results
result.show()
# Save results
result.save('result.jpg')
# 4. Training (optional)
# model.train(data='coco128.yaml', epochs=100, imgsz=640)
""")
Image Segmentation
Object detection marks objects with rectangular boxes, while image segmentation goes down to the pixel level, labeling which object each pixel belongs to.
Semantic Segmentation vs Instance Segmentation
Image segmentation has two major branches: semantic segmentation and instance segmentation.
| Task | Goal | Example | Representative Models |
|---|---|---|---|
| Semantic Segmentation | Classify each pixel (regardless of instance) | Mark all people in red, all cars in blue | FCN、U-Net、DeepLab |
| Instance Segmentation | Classify each pixel and distinguish different instances | Mark person 1 in red, person 2 in orange | Mask R-CNN、YOLOv8-Seg |
| Panoptic Segmentation | Semantic + instance, unified representation | Simultaneously label all 'things' and 'background' | Panoptic FPN、MaskFormer |
A classic example: suppose there are two people and two dogs.
Semantic segmentation says: here there are 'people' and 'dogs' (only categories, not counts).
Instance segmentation says: here there are 'person 1', 'person 2', 'dog 1', 'dog 2' (each individual is distinct).
Segment Anything Model(SAM)
In 2023, Meta AI released SAM (Segment Anything Model), which turned segmentation into a prompt-driven interactive task.
SAM's input can be: a point, a box, a piece of text, or a rough mask.
SAM's output is: the precise segmentation mask of the corresponding object.
The design philosophy of SAM is:Promptable, generalizable, zero-shot。
Example
# SAM Segmentation Concept Demo
# Show how to segment using prompts such as points, boxes, etc.
# ============================================
class SimpleSAM:
"""Simplified SAM concept demo"""
def __init__(self):
self.image_embedding = None
print("SAM concept demo initialization")
def set_image(self, image: np.ndarray):
"""Encode image (only once)"""
print("Encoding image to embedding...")
self.image_embedding = "image_encoded"
return self
def predict_from_point(self, point: tuple, point_label: int = 1):
"""
Segment from a point
Parameters:
point: (x, y) coordinates
point_label: 1 for foreground, 0 for background
"""
print(f"Segment from point {point} ({'foreground' if point_label == 1 else 'background'})")
return "mask_from_point"
def predict_from_box(self, box: tuple):
"""Segment from a bounding box"""
x1, y1, x2, y2 = box
print(f"Segment from box [{x1}, {y1}, {x2}, {y2}]")
return "mask_from_box"
def predict_from_points(self, points: list, labels: list):
"""Segment from multiple points"""
print(f"Segment from {len(points)} points")
return "mask_from_points"
# Demonstrate SAM usage workflow
print("=" * 60)
print("EXAMPLE SAM Segmentation Concept Demo")
print("=" * 60)
sam = SimpleSAM()
# 1. Set up image (encode once)
test_img = create_test_image()
sam.set_image(test_img)
# 2. Segment with different prompts
print()
mask1 = sam.predict_from_point((100, 200), point_label=1)
print()
mask2 = sam.predict_from_box((50, 50, 200, 300))
print()
mask3 = sam.predict_from_points([(100, 100), (150, 150)], [1, 1])
# ============================================
# Code example for actual SAM usage
# ============================================
print("\n" + "=" * 60)
print("Code example for actual SAM usage:")
print("=" * 60)
print("""
# Install: pip install segment-anything torch torchvision opencv-python
import cv2
import numpy as np
from segment_anything import sam_model_registry, SamPredictor
# 1. Load model
# Download checkpoint: https://dl.fbaipublicfiles.com/segment_anything/sam_vit_h_4b8939.pth
sam = sam_model_registry["vit_h"](checkpoint="sam_vit_h_4b8939.pth")
sam.to(device="cuda")
predictor = SamPredictor(sam)
# 2. Set up image
image = cv2.imread("test_image.jpg")
image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)
predictor.set_image(image)
# 3. Segment with point prompts
input_point = np.array([[300, 250]]) # prompt point coordinates
input_label = np.array([1]) # 1 = foreground, 0 = background
masks, scores, logits = predictor.predict(
point_coords=input_point,
point_labels=input_label,
multimask_output=True
)
# 4. Visualize results
for i, mask in enumerate(masks):
color = np.array([30, 144, 255])
h, w = mask.shape
mask_image = mask.reshape(h, w, 1) * color.reshape(1, 1, -1)
cv2.imwrite(f"mask_{i}.jpg", mask_image)
""")
Applications of Segmentation in Industry
Image segmentation is a core technology in many industrial applications:
| Field | Application scenario | Task type |
|---|---|---|
| Medical imaging | Tumor segmentation, organ measurement | Semantic segmentation |
| Autonomous driving | Drivable area, lane line detection | Semantic segmentation |
| Remote sensing imagery | Land cover classification, building extraction | Semantic segmentation |
| Intelligent manufacturing | Defect detection, part measurement | Instance segmentation |
| Video editing | Person matting, background replacement | Instance segmentation |
Diffusion Models
Diffusion models are the core technology for AI image generation; Stable Diffusion, DALL-E 3, and Midjourney are all based on it.
Forward Diffusion Process (Adding Noise)
Forward diffusion is a simple process:Gradually add Gaussian noise to the image until it becomes pure noise.。
This is a deterministic process; each step adds only a little noise, and we can use mathematical formulas to precisely calculate the amount of noise at any time.
Example
# Demonstrate the forward process of the diffusion model (adding noise)
# ============================================
import numpy as np
import matplotlib.pyplot as plt
class ForwardDiffusion:
"""Forward diffusion process demo"""
def __init__(self, num_timesteps: int = 1000):
self.num_timesteps = num_timesteps
# Define beta schedule (linearly increasing from beta_start to beta_end)
self.beta_start = 0.0001
self.beta_end = 0.02
self.betas = np.linspace(self.beta_start, self.beta_end, num_timesteps)
# Compute alpha and alpha_bar
self.alphas = 1.0 - self.betas
self.alphas_bar = np.cumprod(self.alphas)
def q_sample(self, x0: np.ndarray, t: int) -> tuple:
"""
Compute xt at time t directly from x0
Parameters:
x0: original image (H, W, 3)
t: time step (0 <= t < num_timesteps)
Returns:
xt: image at time t
noise: the added noise
"""
sqrt_alpha_bar_t = np.sqrt(self.alphas_bar[t])
sqrt_one_minus_alpha_bar_t = np.sqrt(1 - self.alphas_bar[t])
# Generate random noise of the same shape
noise = np.random.randn(*x0.shape)
# Compute xt in one step
xt = sqrt_alpha_bar_t * x0 + sqrt_one_minus_alpha_bar_t * noise
return xt, noise
# Create a simple test image: a bright circle in the middle
def create_diffusion_test_image():
h, w = 64, 64
x = np.linspace(-1, 1, w)
y = np.linspace(-1, 1, h)
xx, yy = np.meshgrid(x, y)
r = np.sqrt(xx**2 + yy**2)
image = np.exp(-r**2 / 0.1) # A bright circle with Gaussian distribution
image = (image - image.min()) / (image.max() - image.min()) # Normalize to [0, 1]
image = np.stack([image, image*0.8, image*0.5], axis=-1) # Convert to 3 channels
return image
print("=" * 60)
print("EXAMPLE Forward Diffusion Process Demo")
print("=" * 60)
# Initialize diffusion process
diffusion = ForwardDiffusion(num_timesteps=1000)
# Create test image
x0 = create_diffusion_test_image()
print(f"Original image shape: {x0.shape}")
# Display the noise-adding results at different time steps
timesteps_to_show = [0, 100, 200, 300, 500, 1000]
print("\nNoise level at different time steps: ")
for t in timesteps_to_show:
if t < 1000:
xt, noise = diffusion.q_sample(x0, t)
alpha_bar = diffusion.alphas_bar[t]
print(f" t={t:4d}: alpha_bar={alpha_bar:.4f}, SNR={alpha_bar/(1-alpha_bar):.4f}")
else:
print(f" t={t:4d}: pure noise")
print("\nThe core formula of forward diffusion: ")
print(" q(x_t | x_0) = N(x_t; sqrt(alpha_bar_t) * x0, (1 - alpha_bar_t) * I)")
print("\nThis formula lets us compute x_t at any time t from x0 in one step, without step-by-step iteration!)
Reverse Denoising Process
The reverse process is the inverse of the forward process:Starting from pure noise, gradually predict and remove noise to finally obtain an image。
This is what our model needs to learn: given xt and t, predict what the added noise is.
Training objective: minimize the mean squared error between the noise predicted by the model and the true noise.
U-Net Architecture
The core of a diffusion model is a U-Net network, whose shape resembles the letter "U": downsampling and compressing on the left, upsampling and recovering on the right.
The key to U-Net is the skip connection, which directly passes features from downsampling to the corresponding upsampling layers, preserving detailed information.
DDPM vs DDIM Sampling
After training the model, we need to sample from noise to generate images. Different samplers vary in speed and quality.
| Sampler | Full name | Steps | Characteristics |
|---|---|---|---|
| DDPM | Denoising Diffusion Probabilistic Models | 1000 | Original method, good quality but slow |
| DDIM | Denoising Diffusion Implicit Models | 50-100 | Deterministic sampling, fast |
| LMS | Linear Multistep Method | 30-50 | Fast, good quality |
| DPM-Solver | DPM-Solver | 10-20 | Ultra-fast, achieves good results in just a few steps |
Current Stable Diffusion uses DPM-Solver++ by default, usually achieving high-quality results in 20-30 steps, tens of times faster than the original DDPM.
Stable Diffusion Architecture
Stable Diffusion does not work directly in pixel space, but in latent space (Latent Space), which makes it faster and more efficient.
The complete Stable Diffusion consists of three parts:
1. VAE (Variational Autoencoder): compresses images into latent vectors, or reconstructs images from latent vectors.
2. UNet: the core denoising network, takes latent vectors, time steps, and prompt embeddings, and predicts noise.
3. Text Encoder (CLIP): converts text prompts into embeddings, telling the model what to generate.
Example
# Stable Diffusion concept demonstration
# Show the complete generation pipeline
# ============================================
class SimpleStableDiffusion:
"""Simplified Stable Diffusion concept demonstration"""
def __init__(self):
print("Stable Diffusion concept demonstration initialization")
def encode_text(self, prompt: str):
"""Encode text prompt"""
print(f"Encode prompt: '{prompt}'")
return "text_embedding"
def encode_image(self, image: np.ndarray):
"""VAE encodes image to latent space"""
print(Encode image to latent space (VAE encoder))
return "latent_code"
def decode_latent(self, latent: str):
VAE decodes image from latent space
print(Decode image from latent space (VAE decoder))
return "decoded_image"
def denoise_step(self, latent: str, text_embedding: str, t: int):
One-step denoising
print(fDenoising step t={t})
return "denoised_latent"
def generate(self, prompt: str, num_inference_steps: int = 50):
Complete generation pipeline
print("=" * 60)
print(fStarting generation: '{prompt}')
print("=" * 60)
# 1. Encode text
text_emb = self.encode_text(prompt)
# 2. Start from random noise
print("\nInitializing random noise...)
latent = "random_noise"
# 3. Denoise step by step
print(f"\nStarting denoising loop ({num_inference_steps} steps):)
for t in reversed(range(num_inference_steps)):
latent = self.denoise_step(latent, text_emb, t)
# 4. Decode to obtain image
print("\nDecoding latent vector to image...)
image = self.decode_latent(latent)
print("\nGeneration complete!)
return image
# Demonstrate generation pipeline
sd = SimpleStableDiffusion()
image = sd.generate(
prompt=A cute orange cat in the garden,
num_inference_steps=20
)
# ============================================
# Actual code example using Stable Diffusion
# ============================================
print("\n" + "=" * 60)
print(Actual code example using Stable Diffusion:)
print("=" * 60)
print("""
# Installation: pip install diffusers transformers accelerate torch
from diffusers import StableDiffusionPipeline
import torch
# 1. Load the model
model_id = "runwayml/stable-diffusion-v1-5"
pipe = StableDiffusionPipeline.from_pretrained(
model_id,
torch_dtype=torch.float16
)
pipe = pipe.to("cuda")
# 2. Text-to-image
prompt = "a cute orange cat in a garden"
image = pipe(prompt).images[0]
# 3. Save the result
image.save("cat_garden.png")
# 4. Image-to-image
from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image
img_pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
model_id,
torch_dtype=torch.float16
).to("cuda")
init_image = Image.open("input.jpg").convert("RGB")
init_image = init_image.resize((512, 512))
image = img_pipe(
prompt="make it look like a painting",
image=init_image,
strength=0.75 # degree of variation
).images[0]
image.save("painted.png")
# 5. Inpainting (local modification)
from diffusers import StableDiffusionInpaintPipeline
inpaint_pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_id,
torch_dtype=torch.float16
).to("cuda")
image = inpaint_pipe(
prompt="put a hat on the cat",
image=init_image,
mask_image=mask_image # mask image, white areas are the parts to be modified
).images[0]
""")
CLIP: Foundations of Multimodal Alignment
CLIP (Contrastive Language-Image Pre-training) is a work published by OpenAI in 2021, which for the first time truly aligned text and images into the same vector space.
Principles of Contrastive Learning
CLIP's training method is contrastive learning:Make correct image-text pairs have high similarity, and incorrect pairs have low similarity.。
During training, a batch contains N images and N text descriptions (in one-to-one correspondence), forming N×N possible image-text pairs.
The diagonal contains matching positive samples, and the rest are unmatched negative samples.
The training objective is: maximize the similarity of positive samples and minimize the similarity of negative samples.
Image-Text Alignment Training
CLIP consists of two encoders:
1. Image Encoder: encodes an image into a vector (can be ResNet or ViT).
-
2. Text Encoder: encodes text into vectors (Transformer).
After the two vectors are projected into the same space, the degree of image-text matching is computed via cosine similarity.
Example
# CLIP concept demo: contrastive learning and image-text alignment
# ============================================
import torch
import torch.nn.functional as F
class SimpleCLIP:
"""Simplified CLIP concept demo"""
def __init__(self, embed_dim: int = 512):
self.embed_dim = embed_dim
# Simple simulation: not real encoding, just random projection
print("CLIP concept demo initialization")
print(f"Embedding dimension: {embed_dim}")
def encode_image(self, image: str) -> torch.Tensor:
"""Simulate image encoding"""
# In practice, this would be a Vision Transformer or ResNet
return torch.randn(1, self.embed_dim)
def encode_text(self, text: str) -> torch.Tensor:
"""Simulate text encoding"""
# In practice, this would be a Text Transformer
return torch.randn(1, self.embed_dim)
def compute_similarity(self, image_emb: torch.Tensor, text_emb: torch.Tensor) -> float:
"""Compute cosine similarity"""
image_emb = F.normalize(image_emb, dim=-1)
text_emb = F.normalize(text_emb, dim=-1)
return (image_emb @ text_emb.T).item()
# Demonstrate CLIP inference flow
print("=" * 60)
print("EXAMPLE CLIP zero-shot classification demo")
print("=" * 60)
clip = SimpleCLIP(embed_dim=512)
# Assume there is an image
print("\n1. Encoding image...")
image_emb = clip.encode_image("A photo of a cat")
# Assume there are several candidate categories
candidates = ["A photo of a dog", "A photo of a cat", "A car", "A person"]
print(f"\n2. Encoding candidate texts: {candidates}")
# Compute similarities
print("\n3. Computing image-text similarity:")
for text in candidates:
text_emb = clip.encode_text(text)
sim = clip.compute_similarity(image_emb, text_emb)
print(f" '{text}' → similarity: {sim:.3f}")
print("\n(In practice, 'a photo of a cat' will have the highest similarity)")
# ============================================
# Actual example of using CLIP
# ============================================
print("\n" + "=" * 60)
print("Actual example of using CLIP:")
print("=" * 60)
print("""
# Install: pip install clip torch torchvision
import clip
import torch
from PIL import Image
# 1. Load the model
device = "cuda" if torch.cuda.is_available() else "cpu"
model, preprocess = clip.load("ViT-B/32", device=device)
# 2. Prepare the image
image = preprocess(Image.open("cat.jpg")).unsqueeze(0).to(device)
# 3. Prepare the text
text = clip.tokenize([
"a photo of a dog",
"a photo of a cat",
"a photo of a car",
"a photo of a person"
]).to(device)
# 4. Inference
with torch.no_grad():
image_features = model.encode_image(image)
text_features = model.encode_text(text)
logits_per_image, logits_per_text = model(image, text)
probs = logits_per_image.softmax(dim=-1).cpu().numpy()
# 5. Output the results
print("Prediction probabilities:")
for i, p in enumerate(probs[0]):
print(f" {i+1}. {p:.2%}")
""")
Zero-Shot Classification Capability
The most exciting capability of CLIP is Zero-Shot classification:It can classify directly without training on the target dataset。
Models trained on traditional ImageNet can only classify the 1000 classes of ImageNet; adding new classes requires retraining.
CLIP can use arbitrary text to describe categories, for example: "a photo of a cat", "an orange cat sitting on the grass", "a sketch of a cat".
The significance of CLIP is not only zero-shot classification, but also that it establishes a general "vision-language" interface; many subsequent models (such as Stable Diffusion, vision-language models) are built on its foundation.
Vision-Language Models (VLM)
A Vision-Language Model (VLM) combines image understanding and language generation, enabling the model to "look at pictures and talk".
LLaVA Architecture
LLaVA (Large Language and Vision Assistant) is an open-source VLM. Its design is simple and clear, making it very suitable for beginners.
The architecture of LLaVA is divided into three parts:
-
1. Vision Encoder: CLIP's vision encoder, converting images into embeddings.
-
2. Projection Layer: projects the visual embeddings into the same dimension as the LLM input.
-
3. LLM: a large language model (e.g., Vicuna, LLaMA) that receives the concatenated image-text input and generates answers.
Training is divided into two steps:
-
1. Pretraining phase: only the Projection Layer is trained, using a large number of image-text pairs for alignment.
-
2. Instruction fine-tuning phase: the entire model is trained with visual instruction datasets (e.g., LLaVA-Instruct).
Connection between Visual Encoder and LLM
The key to connecting vision and language lies in:Representing images as token sequences that language models can understand。
There are several common connection methods:
| Method | Characteristics | Representative models |
|---|---|---|
| Special token | Use one [IMG] token to represent the entire image | Early models |
| Sequence concatenation | Directly concatenate image patch embeddings before the text | LLaVA、GPT-4V |
| Cross attention | Add cross-attention to visual features in the LLM | BLIP-2、Flamingo |
Example
# Vision-Language Model (VLM) Concept Demo
# Show how to concatenate visual and language inputs
# ============================================
class SimpleVLM:
"""Simplified vision-language model"""
def __init__(self, vocab_size: int = 50000, hidden_size: int = 768):
self.vocab_size = vocab_size
self.hidden_size = hidden_size
print("Simplified VLM initialization")
def encode_vision(self, image: str) -> torch.Tensor:
"""Encode image as sequence"""
# Assume there are 256 visual tokens
return torch.randn(1, 256, self.hidden_size)
def encode_text(self, text: str) -> torch.Tensor:
"""Encode text as sequence"""
# Assume the text has 16 tokens
return torch.randn(1, 16, self.hidden_size)
def generate(self, image: str, question: str) -> str:
"""Image-text concatenation generates answer"""
print("=" * 60)
print(f"Question: {question}")
print("=" * 60)
# 1. Encode image and text
print("\n1. Encode image...)
vision_emb = self.encode_vision(image)
print("2. Encode text...")
text_emb = self.encode_text(question)
# 3. Concatenate: [vision sequence] + [text sequence]
print("\n3. Concatenate vision and text sequences...)
print(f" Vision embedding shape: {vision_emb.shape}")
print(f" Text embedding shape: {text_emb.shape}")
combined = torch.cat([vision_emb, text_emb], dim=1)
print(f" Shape after concatenation: {combined.shape}")
# 4. LLM generates answer
print("\n4. LLM generates answer...)
return "This is a cute orange cat, it looks very happy."
# Demonstrate VLM Q&A
vlm = SimpleVLM()
answer = vlm.generate(
image="A photo of an orange cat",
question="What is in this picture? Please describe it."
)
print(f"\nAnswer: {answer}")
# ============================================
# Code example for actually using LLaVA
# ============================================
print("\n" + "=" * 60)
print("Code example for actually using LLaVA:")
print("=" * 60)
print("""
# LLaVA requires a lot of GPU memory, it is recommended to use the transformers or llava library
from transformers import AutoProcessor, AutoModelForVisionAndLanguageGeneration
import torch
from PIL import Image
# 1. Load model and processor
model_id = "llava-hf/llava-1.5-7b-hf"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForVisionAndLanguageGeneration.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
# 2. Prepare input
image = Image.open("cat.jpg")
prompt = "USER: <image>\\nWhat is in this picture? Please describe it. ASSISTANT:"
inputs = processor(images=image, text=prompt, return_tensors="pt").to("cuda")
# 3. Generate
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=500)
answer = processor.decode(output[0], skip_special_tokens=True)
print(answer)
""")
Hands-On: Complete Project for Image Classification + Object Detection
Now let's put together the knowledge we've learned to build a complete hands-on project.
Example
# Hands-on Project: Complete Vision AI Pipeline
# Includes image classification, object detection, and result visualization
# ============================================
import torch
import torch.nn as nn
import torch.nn.functional as F
from torch.utils.data import Dataset, DataLoader
import numpy as np
from PIL import Image
import json
from pathlib import Path
class VisionPipeline:
"""
Complete Vision AI Pipeline
Features:
1. Image classification
2. Object detection
3. Result summary and visualization
"""
def __init__(self, device: str = "cpu"):
self.device = device
self.classifier = None
self.detector = None
print(f"EXAMPLE Vision Pipeline initialized (device: {device})")
def build_classifier(self, num_classes: int = 10):
"""Build image classifier"""
print("\nBuilding image classifier...)
# Use the previously defined SimpleCNN
self.classifier = SimpleCNN(num_classes=num_classes).to(self.device)
print("Image classifier built successfully!")
return self
def build_detector(self):
"""Build object detector"""
print("\nBuilding object detector...)
self.detector = SimpleYOLO()
print("Detector built successfully!")
return self
def load_pretrained_weights(self, classifier_path: str = None):
"""Load pretrained weights"""
if classifier_path and self.classifier:
print(f"\nLoad classifier weights: {classifier_path}")
# In practice: self.classifier.load_state_dict(torch.load(classifier_path))
return self
def classify_image(self, image: np.ndarray) -> tuple:
"""
Image Classification
Returns:
(pred_class_id, pred_class_name, confidence)
"""
if self.classifier is None:
raise ValueError("Please build the classifier first: build_classifier()")
# Preprocessing
if isinstance(image, np.ndarray):
if image.ndim == 2: # Grayscale
image = np.stack([image]*3, axis=-1)
image = torch.from_numpy(image).permute(2, 0, 1).float() / 255.0
image = F.interpolate(image.unsqueeze(0), size=(32, 32), mode='bilinear')
image = image.to(self.device)
# Inference
self.classifier.eval()
with torch.no_grad():
logits = self.classifier(image)
probs = F.softmax(logits, dim=-1)
confidence, pred_idx = probs.max(dim=-1)
# Return result
class_names = [
'airplane', 'automobile', 'bird', 'cat', 'deer',
'dog', 'frog', 'horse', 'ship', 'truck'
]
pred_class = class_names[pred_idx.item()]
return pred_idx.item(), pred_class, confidence.item()
def detect_objects(self, image: np.ndarray) -> list:
"""Object Detection"""
if self.detector is None:
raise ValueError("Please build the detector first: build_detector()")
# Use simulated detection
return self.detector.predict_dummy(image)
def process_image(self, image: np.ndarray) -> dict:
"""
Process an entire image
Returns:
Dictionary containing all results
"""
print("\n" + "=" * 60)
print("Start processing image")
print("=" * 60)
results = {}
# 1. Image classification
if self.classifier is not None:
print("\n[1/3] Running image classification...")
class_id, class_name, conf = self.classify_image(image)
results['classification'] = {
'class_id': class_id,
'class_name': class_name,
'confidence': conf
}
print(f" → Classification result: {class_name} (Confidence: {conf:.2%})")
# 2. Object detection
if self.detector is not None:
print("\n[2/3] Running object detection...")
detections = self.detect_objects(image)
results['detections'] = detections
print(f" → Detected {len(detections)} objects:")
for det in detections:
print(f" - {det['class_name']} ({det['score']:.2f})")
# 3. Visualization
print("\n[3/3] Generating visualization...")
if self.detector is not None:
vis_image = self.detector.visualize(image, detections)
results['visualization'] = vis_image
print(" → Visualization complete")
print("\n" + "=" * 60)
print("Image processing complete!")
print("=" * 60)
return results
def save_results(self, results: dict, output_dir: str = "output"):
"""Save results to file"""
output_path = Path(output_dir)
output_path.mkdir(exist_ok=True)
# Save JSON results
json_result = {
'classification': results.get('classification'),
'detections': [
{k: v for k, v in det.items() if k != 'bbox'}
for det in results.get('detections', [])
]
}
with open(output_path / "results.json", "w") as f:
json.dump(json_result, f, indent=2, ensure_ascii=False)
# Save visualization image
if 'visualization' in results:
vis = Image.fromarray(results['visualization'])
vis.save(output_path / "visualization.jpg")
print(f"\nResults saved to: {output_path}/")
# Complete training example
class SimpleDataset(Dataset):
"""Simple simulated dataset"""
def __init__(self, num_samples: int = 100):
self.num_samples = num_samples
# Generate fake data
self.images = torch.randn(num_samples, 3, 32, 32)
self.labels = torch.randint(0, 10, (num_samples,))
def __len__(self):
return self.num_samples
def __getitem__(self, idx):
return self.images[idx], self.labels[idx]
def train_classifier(pipeline: VisionPipeline, num_epochs: int = 5):
"""Complete workflow for training the classifier"""
print("\n" + "=" * 60)
print("Start training classifier")
print("=" * 60)
# 1. Prepare data
print("\nPreparing dataset...")
train_dataset = SimpleDataset(1000)
val_dataset = SimpleDataset(200)
train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True)
val_loader = DataLoader(val_dataset, batch_size=32)
print(f"Training set: {len(train_dataset)} samples")
print(f"Validation set: {len(val_dataset)} samples")
# 2. Optimizer and loss function
optimizer = torch.optim.Adam(pipeline.classifier.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
# 3. Training loop
print("\nStart training...")
for epoch in range(num_epochs):
pipeline.classifier.train()
total_loss = 0.0
correct = 0
total = 0
for images, labels in train_loader:
images, labels = images.to(pipeline.device), labels.to(pipeline.device)
optimizer.zero_grad()
outputs = pipeline.classifier(images)
loss = criterion(outputs, labels)
loss.backward()
optimizer.step()
total_loss += loss.item()
_, predicted = outputs.max(1)
total += labels.size(0)
correct += predicted.eq(labels).sum().item()
train_acc = correct / total
avg_loss = total_loss / len(train_loader)
# Validation
pipeline.classifier.eval()
correct = 0
total = 0
with torch.no_grad():
for images, labels in val_loader:
images, labels = images.to(pipeline.device), labels.to(pipeline.device)
outputs = pipeline.classifier(images)
_, predicted = outputs.max(1)
total += labels.size(0)
correct += predicted.eq(labels).sum().item()
val_acc = correct / total
print(f"Epoch {epoch+1}/{num_epochs} - "
f"Loss: {avg_loss:.4f}, "
f"Training accuracy: {train_acc:.2%}, "
fValidation accuracy: {val_acc:.2%})
print("\nTraining complete!)
# ============================================
Run the complete hands-on project.
# ============================================
print("=" * 60)
print(EXAMPLE Visual AI Complete Practical Project)
print("=" * 60)
# 1. Create Pipeline
pipeline = VisionPipeline(device="cuda" if torch.cuda.is_available() else "cpu")
# 2. Build the Model
pipeline.build_classifier(num_classes=10)
pipeline.build_detector()
3. Training the classifier
train_classifier(pipeline, num_epochs=3)
# 4. Processing Images
test_image = create_test_image()
results = pipeline.process_image(test_image)
5. Save the result
pipeline.save_results(results, output_dir="example_output")
print("\n" + "=" * 60)
print(Project run completed!)
print("=" * 60)
print("\nNext step learning suggestions:)
print(1. Try training on a real dataset (e.g., CIFAR-10).)
print(2. Fine-tuning using a pretrained model)
print(3. Try more complex model architectures.)
print(4. Deploy to practical applications.)