Computer Vision AI

Of the information humans acquire, about 80% comes from vision.

When you see a photo, you can immediately recognize how many people are in it, what they are doing, and whether the background is indoors or outdoors.

But for a computer, this photo is just a pile of numbers—each pixel is represented by three values for red, green, and blue, and that's all.

Computer Vision (CV) is the technology that enables computers to understand images.

From facial unlock on phones to road condition recognition in autonomous driving, and to lesion analysis in medical imaging, computer vision has become ubiquitous.

This module will take you from basic convolutional neural networks all the way to the latest vision-language models.

Learning path: Convolutional Neural Networks → Vision Transformer → Object Detection → Image Segmentation → Diffusion Models → CLIP → Vision-Language Models. Each step comes with runnable code examples.


Convolutional Neural Networks (CNN)

CNN is a foundational technology in computer vision. It mimics the way the human visual cortex works, extracting image features through local receptive fields.

Principles of Convolution Operation

The core idea of convolution is:Use a small sliding window (convolution kernel) to scan the image and extract local features.。

For example, a 3×3 kernel looks at 9 pixels at a time, computes their weighted sum, and produces an output value.

This process is repeated; the kernel slides over the image from left to right and top to bottom, ultimately generating a "feature map".

Example

# ============================================
# Implement the simplest convolution operation using NumPy
# Demonstrate how convolution extracts edge features
# ============================================

import numpy as np


def simple_convolution(image: np.ndarray, kernel: np.ndarray) -> np.ndarray:
    """
Implement the most basic 2D convolution operation (without padding and stride)

Parameters:
image: input image (H, W), single-channel grayscale
kernel: convolution kernel (kH, kW)

Returns:
Feature map after convolution
    """

    # Get the dimensions of the image and the kernel
    img_h, img_w = image.shape
    kernel_h, kernel_w = kernel.shape

    # Compute the size of the output feature map
    # Output size = input size - kernel size + 1
    out_h = img_h - kernel_h + 1
    out_w = img_w - kernel_w + 1

    # Initialize the output feature map
    output = np.zeros((out_h, out_w))

    # Slide the convolution kernel to compute
    for i in range(out_h):
        for j in range(out_w):
            # Extract the local region of the image corresponding to the kernel
            region = image[i:i+kernel_h, j:j+kernel_w]
            # Multiply corresponding elements and sum (this is the convolution operation)
            output[i, j] = np.sum(region * kernel)

    return output


# Create a simple test image: a white square in the middle, black around it
# Shape: 8×8 grayscale image
test_image = np.array([
    [0, 0, 0, 0, 0, 0, 0, 0],
    [0, 0, 0, 0, 0, 0, 0, 0],
    [0, 0, 1, 1, 1, 1, 0, 0],
    [0, 0, 1, 1, 1, 1, 0, 0],
    [0, 0, 1, 1, 1, 1, 0, 0],
    [0, 0, 1, 1, 1, 1, 0, 0],
    [0, 0, 0, 0, 0, 0, 0, 0],
    [0, 0, 0, 0, 0, 0, 0, 0],
])
print("Original image:")
print(test_image)

# Define an edge detection kernel (simplified Sobel operator)
# This kernel can detect vertical edges
edge_kernel = np.array([
    [-1, 0, 1],
    [-2, 0, 2],
    [-1, 0, 1],
])

# Perform convolution
feature_map = simple_convolution(test_image, edge_kernel)
print("\nFeature map after convolution (vertical edges detected):")
print(np.round(feature_map, 2))

The key to the convolution operation is:The parameters of the convolution kernel are learned, not manually designed.。

During training, the model automatically adjusts the kernel values, allowing it to extract features useful for the task.

Pooling

Pooling compresses the size of the feature map, reduces computation, while preserving important features.

The most commonly used is max pooling: divide the feature map into several small blocks, and each block retains only the maximum value.

Example

# ============================================
# Implement the max pooling operation
# ============================================

def max_pooling(feature_map: np.ndarray, pool_size: int = 2) -> np.ndarray:
    """
Implement the max pooling operation

Parameters:
feature_map: input feature map (H, W)
pool_size: pooling window size, default 2×2

Returns:
The pooled feature map
    """

    h, w = feature_map.shape
    # Calculate output size
    out_h = h // pool_size
    out_w = w // pool_size

    output = np.zeros((out_h, out_w))

    for i in range(out_h):
        for j in range(out_w):
            # Extract the pooling window region
            region = feature_map[
                i*pool_size:(i+1)*pool_size,
                j*pool_size:(j+1)*pool_size
            ]
            # Take the maximum value
            output[i, j] = np.max(region)

    return output


# Perform pooling using the feature map from earlier
print("Feature map before pooling:")
print(feature_map)

pooled = max_pooling(feature_map, pool_size=2)
print("\n"Feature map after 2×2 max pooling:")
print(pooled)

Classic Architecture Comparison

In the history of CNN development, there are several milestone architectures, each representing design ideas from different periods.

ArchitectureYearCore innovationCharacteristics
AlexNet2012ReLU activation, Dropout, data augmentationThe pioneer of the deep learning era, first to significantly outperform traditional methods on ImageNet.
VGG2014Uniform use of 3×3 small convolution kernelsSimple and elegant structure, easy to understand and implement
ResNet2015Residual connection (Skip Connection)Solves the vanishing gradient problem in deep networks, allowing networks to have hundreds of layers.
EfficientNet2019Compound scaling methodScales depth, width, and resolution simultaneously, with extremely high parameter efficiency.

Residual Connection (Skip Connection)

The core innovation of ResNet is the residual connection, which solves the problem that "the deeper the network, the harder it is to train."

The learning goal of a traditional network layer is to directly learn the mapping from input x to output y.

The learning goal of a residual network is:Learn the difference (residual) between output y and input x.。

Formula: y = F(x) + x, where F(x) is the residual to be learned.

Intuition of residual connection: letting the network learn "how much to change based on the existing state" is much easier than letting it "learn the complete mapping from scratch."

Example

# ============================================
# Implement a complete CNN image classifier with PyTorch
# Includes convolution, pooling, residual connections
# ============================================

import torch
import torch.nn as nn
import torch.nn.functional as F


class ResidualBlock(nn.Module):
    """A simple residual block"""

    def __init__(self, in_channels: int, out_channels: int, stride: int = 1):
        super().__init__()
        # First convolutional layer
        self.conv1 = nn.Conv2d(
            in_channels, out_channels,
            kernel_size=3, stride=stride, padding=1, bias=False
        )
        self.bn1 = nn.BatchNorm2d(out_channels)

        # Second convolutional layer
        self.conv2 = nn.Conv2d(
            out_channels, out_channels,
            kernel_size=3, stride=1, padding=1, bias=False
        )
        self.bn2 = nn.BatchNorm2d(out_channels)

        # Shortcut connection (for matching dimensional changes)
        self.shortcut = nn.Sequential()
        if stride != 1 or in_channels != out_channels:
            self.shortcut = nn.Sequential(
                nn.Conv2d(
                    in_channels, out_channels,
                    kernel_size=1, stride=stride, bias=False
                ),
                nn.BatchNorm2d(out_channels)
            )

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        # Main path: conv → BN → ReLU → conv → BN
        out = F.relu(self.bn1(self.conv1(x)))
        out = self.bn2(self.conv2(out))

        # Residual connection: add the input (or the input after dimension transformation)
        out += self.shortcut(x)

        # Final ReLU
        out = F.relu(out)
        return out


class SimpleCNN(nn.Module):
    """Simple CNN for image classification (with residual connections)"""

    def __init__(self, num_classes: int = 10):
        super().__init__()
        # Initial convolutional layer
        self.conv1 = nn.Conv2d(3, 32, kernel_size=3, stride=1, padding=1)
        self.bn1 = nn.BatchNorm2d(32)

        # Residual layer
        self.layer1 = ResidualBlock(32, 32)
        self.layer2 = ResidualBlock(32, 64, stride=2)
        self.layer3 = ResidualBlock(64, 64)

        # Global average pooling
        self.global_pool = nn.AdaptiveAvgPool2d((1, 1))

        # Classification head
        self.fc = nn.Linear(64, num_classes)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        # x shape: (batch_size, 3, 32, 32)
        x = F.relu(self.bn1(self.conv1(x)))

        x = self.layer1(x)  # (batch_size, 32, 32, 32)
        x = self.layer2(x)  # (batch_size, 64, 16, 16)
        x = self.layer3(x)  # (batch_size, 64, 16, 16)

        x = self.global_pool(x)  # (batch_size, 64, 1, 1)
        x = x.view(x.size(0), -1)  # (batch_size, 64)
        x = self.fc(x)  # (batch_size, num_classes)

        return x


# Create the model and test
model = SimpleCNN(num_classes=10)
print("Model structure:")
print(model)

# Test forward pass
# Create a fake batch: 4 images, each 3×32×32 (RGB 32×32)
test_input = torch.randn(4, 3, 32, 32)
test_output = model(test_input)
print(f"\nInput shape: {test_input.shape}")
print(f"Output shape: {test_output.shape} (batch_size, num_classes)")

Vision Transformer(ViT)

In 2020, Google published the paper "An Image is Worth 16x16 Words", bringing the Transformer architecture from the language domain to the vision domain.

Image Patching (Patch Embedding)

Transformers process sequence data, but images are 2D grid data.

ViT's solution is simple:Cut the image into small patches, and treat each patch as a word。

For example, a 224×224 image cut into 16×16 patches yields 14×14 = 196 patches.

Each patch is flattened into a 1D vector and projected into an embedding via a linear layer, which can then be fed into the Transformer.

Example

# ============================================
# Implement Patch Embedding (image patching)
# ============================================

class PatchEmbedding(nn.Module):
    """
Segment the image into patches and embed them

Parameters:
img_size: input image size (square)
patch_size: size of each patch (square)
in_channels: number of input channels (3 for RGB)
embed_dim: output embedding dimension
    """


    def __init__(
        self,
        img_size: int = 224,
        patch_size: int = 16,
        in_channels: int = 3,
        embed_dim: int = 768
    ):
        super().__init__()
        self.img_size = img_size
        self.patch_size = patch_size
        # Calculate the number of patches: (224/16)^2 = 14^2 = 196
        self.num_patches = (img_size // patch_size) ** 2

        # Use convolution to implement patch partitioning and embedding
        # This is equivalent to: cut patches → flatten → linear projection
        self.proj = nn.Conv2d(
            in_channels, embed_dim,
            kernel_size=patch_size, stride=patch_size
        )

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        # x shape: (batch_size, in_channels, img_size, img_size)
        x = self.proj(x)  # (batch_size, embed_dim, num_patches^0.5, num_patches^0.5)
        x = x.flatten(2)  # (batch_size, embed_dim, num_patches)
        x = x.transpose(1, 2)  # (batch_size, num_patches, embed_dim)
        return x


# Test Patch Embedding
patch_embed = PatchEmbedding(img_size=224, patch_size=16, embed_dim=768)
test_image = torch.randn(1, 3, 224, 224)  # 1 RGB image of 224×224
patches = patch_embed(test_image)
print(f"Input image shape: {test_image.shape}")
print(f"Patch sequence shape: {patches.shape} (batch_size, num_patches, embed_dim)")

ViT Full Architecture

The complete ViT architecture also requires adding a class token and positional encoding.

The class token is a learnable vector that serves to "aggregate" information from all patches.

Positional encoding is also learnable; it tells the model "where each patch is located in the image".

Example

# ============================================
# Implement a simplified Vision Transformer
# ============================================

class MultiHeadAttention(nn.Module):
    """Multi-head self-attention"""

    def __init__(self, dim: int, num_heads: int):
        super().__init__()
        self.num_heads = num_heads
        self.head_dim = dim // num_heads
        assert self.head_dim * num_heads == dim, "Dimension must be divisible by the number of heads"

        # Q, K, V projection matrices
        self.qkv = nn.Linear(dim, dim * 3)
        # Output projection
        self.proj = nn.Linear(dim, dim)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        batch_size, seq_len, dim = x.shape

        # Compute Q, K, V
        qkv = self.qkv(x)  # (batch_size, seq_len, dim*3)
        qkv = qkv.reshape(
            batch_size, seq_len, 3, self.num_heads, self.head_dim
        )  # (batch_size, seq_len, 3, num_heads, head_dim)
        q, k, v = qkv.unbind(2)  # Each shape: (batch_size, seq_len, num_heads, head_dim)

        # Rearrange dimensions to compute attention
        q = q.transpose(1, 2)  # (batch_size, num_heads, seq_len, head_dim)
        k = k.transpose(1, 2)
        v = v.transpose(1, 2)

        # Compute attention scores
        attention = q @ k.transpose(-2, -1) / (self.head_dim ** 0.5)
        attention = attention.softmax(dim=-1)

        # Aggregate values
        out = attention @ v  # (batch_size, num_heads, seq_len, head_dim)
        out = out.transpose(1, 2)  # (batch_size, seq_len, num_heads, head_dim)
        out = out.flatten(2)  # (batch_size, seq_len, dim)

        # Output projection
        out = self.proj(out)
        return out


class TransformerBlock(nn.Module):
    """A Transformer block"""

    def __init__(self, dim: int, num_heads: int, mlp_ratio: float = 4.0):
        super().__init__()
        self.norm1 = nn.LayerNorm(dim)
        self.attn = MultiHeadAttention(dim, num_heads)
        self.norm2 = nn.LayerNorm(dim)

        # MLP
        mlp_hidden = int(dim * mlp_ratio)
        self.mlp = nn.Sequential(
            nn.Linear(dim, mlp_hidden),
            nn.GELU(),
            nn.Linear(mlp_hidden, dim)
        )

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        # Residual connection 1
        x = x + self.attn(self.norm1(x))
        # Residual connection 2
        x = x + self.mlp(self.norm2(x))
        return x


class SimpleViT(nn.Module):
    """Simplified Vision Transformer"""

    def __init__(
        self,
        img_size: int = 224,
        patch_size: int = 16,
        in_channels: int = 3,
        num_classes: int = 10,
        embed_dim: int = 192,
        depth: int = 6,
        num_heads: int = 6
    ):
        super().__init__()
        # Patch embedding
        self.patch_embed = PatchEmbedding(
            img_size, patch_size, in_channels, embed_dim
        )
        num_patches = self.patch_embed.num_patches

        # Class token
        self.cls_token = nn.Parameter(torch.zeros(1, 1, embed_dim))

        # Positional encoding
        self.pos_embed = nn.Parameter(
            torch.zeros(1, num_patches + 1, embed_dim)
        )

        # Transformer layers
        self.blocks = nn.ModuleList([
            TransformerBlock(embed_dim, num_heads)
            for _ in range(depth)
        ])

        # Final LayerNorm and classification head
        self.norm = nn.LayerNorm(embed_dim)
        self.head = nn.Linear(embed_dim, num_classes)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        batch_size = x.shape[0]

        # Patch embedding
        x = self.patch_embed(x)  # (batch_size, num_patches, embed_dim)

        # Add class token
        cls_tokens = self.cls_token.expand(batch_size, -1, -1)
        x = torch.cat((cls_tokens, x), dim=1)  # (batch_size, 1+num_patches, embed_dim)

        # Add positional encoding
        x = x + self.pos_embed

        # Pass through Transformer layers
        for block in self.blocks:
            x = block(x)

        # Use only the class token for classification
        x = self.norm(x)
        cls_output = x[:, 0]  # Take the first position (class token)
        x = self.head(cls_output)

        return x


# Create ViT model and test
vit = SimpleViT(
    img_size=32,    # Use smaller images for easier testing
    patch_size=8,   # 8×8 patch
    embed_dim=128,
    depth=4,
    num_heads=4,
    num_classes=10
)
print("Simplified ViT architecture:")
print(vit)

test_input = torch.randn(2, 3, 32, 32)  # 2 32×32 images
test_output = vit(test_input)
print(f"\nInput shape: {test_input.shape}")
print(f"Output shape: {test_output.shape}")

CNN vs ViT Comparison

Two architectures each have their own pros and cons; which one you choose depends on your specific scenario.

CharacteristicsCNNViT
Inductive biasStrong (locality, translation invariance)Weak (learns from data)
Data requirementsWorks with small dataNeeds large amounts of data to show its advantages
Computational efficiencyUsually faster, mature hardware optimizationLarge parameter count, computationally intensive
InterpretabilityEasy to visualize convolution kernelsAttention maps are interpretable, but the overall model is harder to interpret
Long-range dependenciesRequire multiple layers to establishNaturally supports global receptive field
Applicable scenariosLimited data, high real-time requirementsLarge datasets, complex visual tasks

Rule of thumb: use CNN (ResNet, EfficientNet) when data is limited; consider ViT when data reaches millions or more. There are also hybrid architectures (such as ConvNeXt) that combine the advantages of both.


Object Detection

Image classification answers "what is in this image," while object detection answers "what is in this image and where are they located."

The output of object detection is a series of bounding boxes, each corresponding to an object category.

Two-Stage Detection: R-CNN Series

The idea of two-stage detection is:First find regions that may contain objects, then classify these regions。

First stage: generate region proposals, possibly hundreds of boxes indicating "something might be here."

Second stage: extract features from each candidate region, determine what object it is, and adjust the box position.

Evolution of the R-CNN family:

ModelKey improvementsSpeed
R-CNNReplaces traditional features with CNNSlow (tens of seconds per image)
Fast R-CNNROI Pooling, shared feature computationMedium (a few seconds per image)
Faster R-CNNRPN(Region Proposal Network)Fast (hundreds of milliseconds per image)
Mask R-CNNAdds a segmentation branchMedium

One-Stage Detection: YOLO Series

The idea of one-stage detection is:Directly perform dense prediction on the image without generating candidate regions separately。

YOLO (You Only Look Once) is a representative of one-stage detection, extremely fast, and suitable for real-time applications.

YOLO divides the image into an S×S grid, with each grid cell responsible for predicting objects whose center point falls within that cell.

Each grid cell predicts: B bounding boxes (with confidence) and probabilities for C classes.

End-to-End Detection: DETR

DETR (DEtection TRansformer) is a new paradigm proposed by Facebook AI, using Transformer to turn object detection into a direct sequence prediction problem.

Its biggest feature is: no need for non-maximum suppression (NMS), no need for anchor boxes (Anchor), the output is the final result.

Hands-On: YOLOv11 Object Detection

Now we use Ultralytics YOLOv11 to create a complete object detection example.

Example

# ============================================
# YOLOv11 Object Detection Hands-on
# Includes model loading, detection, visualization
# ============================================

import torch
import numpy as np
from PIL import Image, ImageDraw, ImageFont
import matplotlib.pyplot as plt


class SimpleYOLO:
    """Simplified YOLO demo (using pretrained model)"""

    def __init__(self, model_name: str = "yolo11n.pt"):
        """
Initialize YOLO model

Parameters:
model_name: model name, yolo11n.pt (nano, small), yolo11s.pt (small), etc.
        """

        self.model_name = model_name
        self.model = None
        # 80 categories of the COCO dataset
        self.coco_names = [
            'person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus',
            'train', 'truck', 'boat', 'traffic light', 'fire hydrant',
            'stop sign', 'parking meter', 'bench', 'bird', 'cat', 'dog',
            'horse', 'sheep', 'cow', 'elephant', 'bear', 'zebra', 'giraffe',
            'backpack', 'umbrella', 'handbag', 'tie', 'suitcase', 'frisbee',
            'skis', 'snowboard', 'sports ball', 'kite', 'baseball bat',
            'baseball glove', 'skateboard', 'surfboard', 'tennis racket',
            'bottle', 'wine glass', 'cup', 'fork', 'knife', 'spoon', 'bowl',
            'banana', 'apple', 'sandwich', 'orange', 'broccoli', 'carrot',
            'hot dog', 'pizza', 'donut', 'cake', 'chair', 'couch',
            'potted plant', 'bed', 'dining table', 'toilet', 'tv', 'laptop',
            'mouse', 'remote', 'keyboard', 'cell phone', 'microwave', 'oven',
            'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
            'scissors', 'teddy bear', 'hair drier', 'toothbrush'
        ]
        # Color for each category
        self.colors = np.random.randint(0, 255, (80, 3))

    def load_model(self):
        """Load pretrained model (here using a simulated implementation to demonstrate the concept)"""
        print(f"Loading model {self.model_name}...")
        print("Note: In actual use, please install ultralytics and call:")
        print("  from ultralytics import YOLO")
        print("  model = YOLO('yolo11n.pt')")
        self.model = "pretrained_model_loaded"
        print("Model loaded successfully!")
        return self

    def predict_dummy(self, image: np.ndarray) -> list:
        """
Simulated detection results (for demonstration)
In actual use, call model(image)

Returns:
Detection result list, each element is a dictionary:
            {'bbox': [x1, y1, x2, y2], 'class_id': int, 'score': float, 'class_name': str}
        """

        # Here we simulate detecting several objects
        h, w = image.shape[:2]
        dummy_results = [
            {'bbox': [w*0.1, h*0.2, w*0.3, h*0.6], 'class_id': 0, 'score': 0.92, 'class_name': 'person'},
            {'bbox': [w*0.35, h*0.4, w*0.7, h*0.8], 'class_id': 2, 'score': 0.85, 'class_name': 'car'},
            {'bbox': [w*0.75, h*0.3, w*0.9, h*0.55], 'class_id': 16, 'score': 0.78, 'class_name': 'dog'},
        ]
        return dummy_results

    def visualize(self, image: np.ndarray, results: list) -> np.ndarray:
        """
Draw detection boxes and labels on the image

Parameters:
image: original image (H, W, 3)
results: detection result list

Returns:
Annotated image
        """

        if isinstance(image, np.ndarray):
            image = Image.fromarray(image)

        draw = ImageDraw.Draw(image)

        for result in results:
            bbox = result['bbox']
            class_id = result['class_id']
            score = result['score']
            class_name = result['class_name']

            color = tuple(self.colors[class_id % 80].tolist())

            # Draw box
            draw.rectangle(bbox, outline=color, width=3)

            # Label background
            label = f"{class_name} {score:.2f}"
            try:
                font = ImageFont.truetype("arial.ttf", 16)
            except:
                font = ImageFont.load_default()

            # Calculate text size
            bbox_label = draw.textbbox((0, 0), label, font=font)
            text_w = bbox_label[2] - bbox_label[0]
            text_h = bbox_label[3] - bbox_label[1]

            # Draw label background
            label_bg = [bbox[0], bbox[1]-text_h-5, bbox[0]+text_w+5, bbox[1]]
            draw.rectangle(label_bg, fill=color)

            # Draw text
            draw.text((bbox[0]+2, bbox[1]-text_h-3), label, fill=(255,255,255), font=font)

        return np.array(image)


# Create simulated image (for demonstration)
def create_test_image() -> np.ndarray:
    """Create a test image"""
    h, w = 480, 640
    image = np.zeros((h, w, 3), dtype=np.uint8)
    # Gradient background
    image[:, :, 0] = np.linspace(100, 200, w)[np.newaxis, :]
    image[:, :, 1] = np.linspace(150, 100, h)[:, np.newaxis]
    image[:, :, 2] = 180
    return image


# Demonstrate YOLO detection process
print("=" * 60)
print("EXAMPLE YOLOv11 Object Detection Demo")
print("=" * 60)

# Initialize
detector = SimpleYOLO().load_model()

# Create test image
test_image = create_test_image()
print(f"\nInput image size: {test_image.shape}")

# Simulated detection
results = detector.predict_dummy(test_image)
print(f"\n"Detected {len(results)} objects:")
for i, r in enumerate(results, 1):
    print(f" {i}. {r['class_name']} (Confidence: {r['score']:.2f})")

# Visualization
visualized = detector.visualize(test_image, results)
print(f"\nVisualization complete, output image size: {visualized.shape}")


# ============================================
# Actual code for using YOLOv11 (requires installing ultralytics)
# ============================================
print("\n" + "=" * 60)
print("Example code for actually using YOLOv11:")
print("=" * 60)
print("""
# Install: pip install ultralytics

from ultralytics import YOLO

# 1. Load model
model = YOLO('yolo11n.pt') # You can also use 'yolo11s.pt', 'yolo11m.pt', etc.

# 2. Predict
results = model('test_image.jpg') # Can be an image path, video, or 0 (webcam)

# 3. Process results
for result in results:
# Detection boxes
    boxes = result.boxes
# Segmentation masks (if using a segmentation model)
    masks = result.masks
# Display results
    result.show()
# Save results
    result.save('result.jpg')

# 4. Training (optional)
# model.train(data='coco128.yaml', epochs=100, imgsz=640)
"""
)

Image Segmentation

Object detection marks objects with rectangular boxes, while image segmentation goes down to the pixel level, labeling which object each pixel belongs to.

Semantic Segmentation vs Instance Segmentation

Image segmentation has two major branches: semantic segmentation and instance segmentation.

TaskGoalExampleRepresentative Models
Semantic SegmentationClassify each pixel (regardless of instance)Mark all people in red, all cars in blueFCN、U-Net、DeepLab
Instance SegmentationClassify each pixel and distinguish different instancesMark person 1 in red, person 2 in orangeMask R-CNN、YOLOv8-Seg
Panoptic SegmentationSemantic + instance, unified representationSimultaneously label all 'things' and 'background'Panoptic FPN、MaskFormer

A classic example: suppose there are two people and two dogs.

Semantic segmentation says: here there are 'people' and 'dogs' (only categories, not counts).

Instance segmentation says: here there are 'person 1', 'person 2', 'dog 1', 'dog 2' (each individual is distinct).

Segment Anything Model(SAM)

In 2023, Meta AI released SAM (Segment Anything Model), which turned segmentation into a prompt-driven interactive task.

SAM's input can be: a point, a box, a piece of text, or a rough mask.

SAM's output is: the precise segmentation mask of the corresponding object.

The design philosophy of SAM is:Promptable, generalizable, zero-shot。

Example

# ============================================
# SAM Segmentation Concept Demo
# Show how to segment using prompts such as points, boxes, etc.
# ============================================

class SimpleSAM:
    """Simplified SAM concept demo"""

    def __init__(self):
        self.image_embedding = None
        print("SAM concept demo initialization")

    def set_image(self, image: np.ndarray):
        """Encode image (only once)"""
        print("Encoding image to embedding...")
        self.image_embedding = "image_encoded"
        return self

    def predict_from_point(self, point: tuple, point_label: int = 1):
        """
Segment from a point

Parameters:
point: (x, y) coordinates
point_label: 1 for foreground, 0 for background
        """

        print(f"Segment from point {point} ({'foreground' if point_label == 1 else 'background'})")
        return "mask_from_point"

    def predict_from_box(self, box: tuple):
        """Segment from a bounding box"""
        x1, y1, x2, y2 = box
        print(f"Segment from box [{x1}, {y1}, {x2}, {y2}]")
        return "mask_from_box"

    def predict_from_points(self, points: list, labels: list):
        """Segment from multiple points"""
        print(f"Segment from {len(points)} points")
        return "mask_from_points"


# Demonstrate SAM usage workflow
print("=" * 60)
print("EXAMPLE SAM Segmentation Concept Demo")
print("=" * 60)

sam = SimpleSAM()

# 1. Set up image (encode once)
test_img = create_test_image()
sam.set_image(test_img)

# 2. Segment with different prompts
print()
mask1 = sam.predict_from_point((100, 200), point_label=1)

print()
mask2 = sam.predict_from_box((50, 50, 200, 300))

print()
mask3 = sam.predict_from_points([(100, 100), (150, 150)], [1, 1])


# ============================================
# Code example for actual SAM usage
# ============================================
print("\n" + "=" * 60)
print("Code example for actual SAM usage:")
print("=" * 60)
print("""
# Install: pip install segment-anything torch torchvision opencv-python

import cv2
import numpy as np
from segment_anything import sam_model_registry, SamPredictor

# 1. Load model
# Download checkpoint: https://dl.fbaipublicfiles.com/segment_anything/sam_vit_h_4b8939.pth
sam = sam_model_registry["vit_h"](checkpoint="sam_vit_h_4b8939.pth")
sam.to(device="cuda")
predictor = SamPredictor(sam)

# 2. Set up image
image = cv2.imread("test_image.jpg")
image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)
predictor.set_image(image)

# 3. Segment with point prompts
input_point = np.array([[300, 250]]) # prompt point coordinates
input_label = np.array([1]) # 1 = foreground, 0 = background

masks, scores, logits = predictor.predict(
    point_coords=input_point,
    point_labels=input_label,
    multimask_output=True
)

# 4. Visualize results
for i, mask in enumerate(masks):
    color = np.array([30, 144, 255])
    h, w = mask.shape
    mask_image = mask.reshape(h, w, 1) * color.reshape(1, 1, -1)
    cv2.imwrite(f"mask_{i}.jpg", mask_image)
"""
)

Applications of Segmentation in Industry

Image segmentation is a core technology in many industrial applications:

FieldApplication scenarioTask type
Medical imagingTumor segmentation, organ measurementSemantic segmentation
Autonomous drivingDrivable area, lane line detectionSemantic segmentation
Remote sensing imageryLand cover classification, building extractionSemantic segmentation
Intelligent manufacturingDefect detection, part measurementInstance segmentation
Video editingPerson matting, background replacementInstance segmentation

Diffusion Models

Diffusion models are the core technology for AI image generation; Stable Diffusion, DALL-E 3, and Midjourney are all based on it.

Forward Diffusion Process (Adding Noise)

Forward diffusion is a simple process:Gradually add Gaussian noise to the image until it becomes pure noise.。

This is a deterministic process; each step adds only a little noise, and we can use mathematical formulas to precisely calculate the amount of noise at any time.

Example

# ============================================
# Demonstrate the forward process of the diffusion model (adding noise)
# ============================================

import numpy as np
import matplotlib.pyplot as plt


class ForwardDiffusion:
    """Forward diffusion process demo"""

    def __init__(self, num_timesteps: int = 1000):
        self.num_timesteps = num_timesteps

        # Define beta schedule (linearly increasing from beta_start to beta_end)
        self.beta_start = 0.0001
        self.beta_end = 0.02
        self.betas = np.linspace(self.beta_start, self.beta_end, num_timesteps)

        # Compute alpha and alpha_bar
        self.alphas = 1.0 - self.betas
        self.alphas_bar = np.cumprod(self.alphas)

    def q_sample(self, x0: np.ndarray, t: int) -> tuple:
        """
Compute xt at time t directly from x0

Parameters:
x0: original image (H, W, 3)
t: time step (0 <= t < num_timesteps)

Returns:
xt: image at time t
noise: the added noise
        """

        sqrt_alpha_bar_t = np.sqrt(self.alphas_bar[t])
        sqrt_one_minus_alpha_bar_t = np.sqrt(1 - self.alphas_bar[t])

        # Generate random noise of the same shape
        noise = np.random.randn(*x0.shape)

        # Compute xt in one step
        xt = sqrt_alpha_bar_t * x0 + sqrt_one_minus_alpha_bar_t * noise

        return xt, noise


# Create a simple test image: a bright circle in the middle
def create_diffusion_test_image():
    h, w = 64, 64
    x = np.linspace(-1, 1, w)
    y = np.linspace(-1, 1, h)
    xx, yy = np.meshgrid(x, y)
    r = np.sqrt(xx**2 + yy**2)
    image = np.exp(-r**2 / 0.1)  # A bright circle with Gaussian distribution
    image = (image - image.min()) / (image.max() - image.min())  # Normalize to [0, 1]
    image = np.stack([image, image*0.8, image*0.5], axis=-1)  # Convert to 3 channels
    return image


print("=" * 60)
print("EXAMPLE Forward Diffusion Process Demo")
print("=" * 60)

# Initialize diffusion process
diffusion = ForwardDiffusion(num_timesteps=1000)

# Create test image
x0 = create_diffusion_test_image()
print(f"Original image shape: {x0.shape}")

# Display the noise-adding results at different time steps
timesteps_to_show = [0, 100, 200, 300, 500, 1000]

print("\nNoise level at different time steps: ")
for t in timesteps_to_show:
    if t < 1000:
        xt, noise = diffusion.q_sample(x0, t)
        alpha_bar = diffusion.alphas_bar[t]
        print(f" t={t:4d}: alpha_bar={alpha_bar:.4f}, SNR={alpha_bar/(1-alpha_bar):.4f}")
    else:
        print(f" t={t:4d}: pure noise")

print("\nThe core formula of forward diffusion: ")
print("  q(x_t | x_0) = N(x_t; sqrt(alpha_bar_t) * x0, (1 - alpha_bar_t) * I)")
print("\nThis formula lets us compute x_t at any time t from x0 in one step, without step-by-step iteration!)

Reverse Denoising Process

The reverse process is the inverse of the forward process:Starting from pure noise, gradually predict and remove noise to finally obtain an image。

This is what our model needs to learn: given xt and t, predict what the added noise is.

Training objective: minimize the mean squared error between the noise predicted by the model and the true noise.

U-Net Architecture

The core of a diffusion model is a U-Net network, whose shape resembles the letter "U": downsampling and compressing on the left, upsampling and recovering on the right.

The key to U-Net is the skip connection, which directly passes features from downsampling to the corresponding upsampling layers, preserving detailed information.

DDPM vs DDIM Sampling

After training the model, we need to sample from noise to generate images. Different samplers vary in speed and quality.

SamplerFull nameStepsCharacteristics
DDPMDenoising Diffusion Probabilistic Models1000Original method, good quality but slow
DDIMDenoising Diffusion Implicit Models50-100Deterministic sampling, fast
LMSLinear Multistep Method30-50Fast, good quality
DPM-SolverDPM-Solver10-20Ultra-fast, achieves good results in just a few steps

Current Stable Diffusion uses DPM-Solver++ by default, usually achieving high-quality results in 20-30 steps, tens of times faster than the original DDPM.

Stable Diffusion Architecture

Stable Diffusion does not work directly in pixel space, but in latent space (Latent Space), which makes it faster and more efficient.

The complete Stable Diffusion consists of three parts:

1. VAE (Variational Autoencoder): compresses images into latent vectors, or reconstructs images from latent vectors.

2. UNet: the core denoising network, takes latent vectors, time steps, and prompt embeddings, and predicts noise.

3. Text Encoder (CLIP): converts text prompts into embeddings, telling the model what to generate.

Example

# ============================================
# Stable Diffusion concept demonstration
# Show the complete generation pipeline
# ============================================

class SimpleStableDiffusion:
    """Simplified Stable Diffusion concept demonstration"""

    def __init__(self):
        print("Stable Diffusion concept demonstration initialization")

    def encode_text(self, prompt: str):
        """Encode text prompt"""
        print(f"Encode prompt: '{prompt}'")
        return "text_embedding"

    def encode_image(self, image: np.ndarray):
        """VAE encodes image to latent space"""
        print(Encode image to latent space (VAE encoder))
        return "latent_code"

    def decode_latent(self, latent: str):
        VAE decodes image from latent space
        print(Decode image from latent space (VAE decoder))
        return "decoded_image"

    def denoise_step(self, latent: str, text_embedding: str, t: int):
        One-step denoising
        print(fDenoising step t={t})
        return "denoised_latent"

    def generate(self, prompt: str, num_inference_steps: int = 50):
        Complete generation pipeline
        print("=" * 60)
        print(fStarting generation: '{prompt}')
        print("=" * 60)

        # 1. Encode text
        text_emb = self.encode_text(prompt)

        # 2. Start from random noise
        print("\nInitializing random noise...)
        latent = "random_noise"

        # 3. Denoise step by step
        print(f"\nStarting denoising loop ({num_inference_steps} steps):)
        for t in reversed(range(num_inference_steps)):
            latent = self.denoise_step(latent, text_emb, t)

        # 4. Decode to obtain image
        print("\nDecoding latent vector to image...)
        image = self.decode_latent(latent)

        print("\nGeneration complete!)
        return image


# Demonstrate generation pipeline
sd = SimpleStableDiffusion()
image = sd.generate(
    prompt=A cute orange cat in the garden,
    num_inference_steps=20
)


# ============================================
# Actual code example using Stable Diffusion
# ============================================
print("\n" + "=" * 60)
print(Actual code example using Stable Diffusion:)
print("=" * 60)
print("""
# Installation: pip install diffusers transformers accelerate torch

from diffusers import StableDiffusionPipeline
import torch

# 1. Load the model
model_id = "runwayml/stable-diffusion-v1-5"
pipe = StableDiffusionPipeline.from_pretrained(
    model_id,
    torch_dtype=torch.float16
)
pipe = pipe.to("cuda")

# 2. Text-to-image
prompt = "a cute orange cat in a garden"
image = pipe(prompt).images[0]

# 3. Save the result
image.save("cat_garden.png")

# 4. Image-to-image
from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image

img_pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
    model_id,
    torch_dtype=torch.float16
).to("cuda")

init_image = Image.open("input.jpg").convert("RGB")
init_image = init_image.resize((512, 512))

image = img_pipe(
    prompt="make it look like a painting",
    image=init_image,
strength=0.75 # degree of variation
).images[0]
image.save("painted.png")

# 5. Inpainting (local modification)
from diffusers import StableDiffusionInpaintPipeline

inpaint_pipe = StableDiffusionInpaintPipeline.from_pretrained(
    model_id,
    torch_dtype=torch.float16
).to("cuda")

image = inpaint_pipe(
    prompt="put a hat on the cat",
    image=init_image,
mask_image=mask_image # mask image, white areas are the parts to be modified
).images[0]
"""
)

CLIP: Foundations of Multimodal Alignment

CLIP (Contrastive Language-Image Pre-training) is a work published by OpenAI in 2021, which for the first time truly aligned text and images into the same vector space.

Principles of Contrastive Learning

CLIP's training method is contrastive learning:Make correct image-text pairs have high similarity, and incorrect pairs have low similarity.。

During training, a batch contains N images and N text descriptions (in one-to-one correspondence), forming N×N possible image-text pairs.

The diagonal contains matching positive samples, and the rest are unmatched negative samples.

The training objective is: maximize the similarity of positive samples and minimize the similarity of negative samples.

Image-Text Alignment Training

CLIP consists of two encoders:

  • 1. Image Encoder: encodes an image into a vector (can be ResNet or ViT).

  • 2. Text Encoder: encodes text into vectors (Transformer).

After the two vectors are projected into the same space, the degree of image-text matching is computed via cosine similarity.

Example

# ============================================
# CLIP concept demo: contrastive learning and image-text alignment
# ============================================

import torch
import torch.nn.functional as F


class SimpleCLIP:
    """Simplified CLIP concept demo"""

    def __init__(self, embed_dim: int = 512):
        self.embed_dim = embed_dim
        # Simple simulation: not real encoding, just random projection
        print("CLIP concept demo initialization")
        print(f"Embedding dimension: {embed_dim}")

    def encode_image(self, image: str) -> torch.Tensor:
        """Simulate image encoding"""
        # In practice, this would be a Vision Transformer or ResNet
        return torch.randn(1, self.embed_dim)

    def encode_text(self, text: str) -> torch.Tensor:
        """Simulate text encoding"""
        # In practice, this would be a Text Transformer
        return torch.randn(1, self.embed_dim)

    def compute_similarity(self, image_emb: torch.Tensor, text_emb: torch.Tensor) -> float:
        """Compute cosine similarity"""
        image_emb = F.normalize(image_emb, dim=-1)
        text_emb = F.normalize(text_emb, dim=-1)
        return (image_emb @ text_emb.T).item()


# Demonstrate CLIP inference flow
print("=" * 60)
print("EXAMPLE CLIP zero-shot classification demo")
print("=" * 60)

clip = SimpleCLIP(embed_dim=512)

# Assume there is an image
print("\n1. Encoding image...")
image_emb = clip.encode_image("A photo of a cat")

# Assume there are several candidate categories
candidates = ["A photo of a dog", "A photo of a cat", "A car", "A person"]
print(f"\n2. Encoding candidate texts: {candidates}")

# Compute similarities
print("\n3. Computing image-text similarity:")
for text in candidates:
    text_emb = clip.encode_text(text)
    sim = clip.compute_similarity(image_emb, text_emb)
    print(f" '{text}' → similarity: {sim:.3f}")

print("\n(In practice, 'a photo of a cat' will have the highest similarity)")


# ============================================
# Actual example of using CLIP
# ============================================
print("\n" + "=" * 60)
print("Actual example of using CLIP:")
print("=" * 60)
print("""
# Install: pip install clip torch torchvision

import clip
import torch
from PIL import Image

# 1. Load the model
device = "cuda" if torch.cuda.is_available() else "cpu"
model, preprocess = clip.load("ViT-B/32", device=device)

# 2. Prepare the image
image = preprocess(Image.open("cat.jpg")).unsqueeze(0).to(device)

# 3. Prepare the text
text = clip.tokenize([
    "a photo of a dog",
    "a photo of a cat",
    "a photo of a car",
    "a photo of a person"
]).to(device)

# 4. Inference
with torch.no_grad():
    image_features = model.encode_image(image)
    text_features = model.encode_text(text)

    logits_per_image, logits_per_text = model(image, text)
    probs = logits_per_image.softmax(dim=-1).cpu().numpy()

# 5. Output the results
print("Prediction probabilities:")
for i, p in enumerate(probs[0]):
    print(f"  {i+1}. {p:.2%}")
"""
)

Zero-Shot Classification Capability

The most exciting capability of CLIP is Zero-Shot classification:It can classify directly without training on the target dataset。

Models trained on traditional ImageNet can only classify the 1000 classes of ImageNet; adding new classes requires retraining.

CLIP can use arbitrary text to describe categories, for example: "a photo of a cat", "an orange cat sitting on the grass", "a sketch of a cat".

The significance of CLIP is not only zero-shot classification, but also that it establishes a general "vision-language" interface; many subsequent models (such as Stable Diffusion, vision-language models) are built on its foundation.


Vision-Language Models (VLM)

A Vision-Language Model (VLM) combines image understanding and language generation, enabling the model to "look at pictures and talk".

LLaVA Architecture

LLaVA (Large Language and Vision Assistant) is an open-source VLM. Its design is simple and clear, making it very suitable for beginners.

The architecture of LLaVA is divided into three parts:

  • 1. Vision Encoder: CLIP's vision encoder, converting images into embeddings.

  • 2. Projection Layer: projects the visual embeddings into the same dimension as the LLM input.

  • 3. LLM: a large language model (e.g., Vicuna, LLaMA) that receives the concatenated image-text input and generates answers.

Training is divided into two steps:

  • 1. Pretraining phase: only the Projection Layer is trained, using a large number of image-text pairs for alignment.

  • 2. Instruction fine-tuning phase: the entire model is trained with visual instruction datasets (e.g., LLaVA-Instruct).

Connection between Visual Encoder and LLM

The key to connecting vision and language lies in:Representing images as token sequences that language models can understand。

There are several common connection methods:

MethodCharacteristicsRepresentative models
Special tokenUse one [IMG] token to represent the entire imageEarly models
Sequence concatenationDirectly concatenate image patch embeddings before the textLLaVA、GPT-4V
Cross attentionAdd cross-attention to visual features in the LLMBLIP-2、Flamingo

Example

# ============================================
# Vision-Language Model (VLM) Concept Demo
# Show how to concatenate visual and language inputs
# ============================================

class SimpleVLM:
    """Simplified vision-language model"""

    def __init__(self, vocab_size: int = 50000, hidden_size: int = 768):
        self.vocab_size = vocab_size
        self.hidden_size = hidden_size
        print("Simplified VLM initialization")

    def encode_vision(self, image: str) -> torch.Tensor:
        """Encode image as sequence"""
        # Assume there are 256 visual tokens
        return torch.randn(1, 256, self.hidden_size)

    def encode_text(self, text: str) -> torch.Tensor:
        """Encode text as sequence"""
        # Assume the text has 16 tokens
        return torch.randn(1, 16, self.hidden_size)

    def generate(self, image: str, question: str) -> str:
        """Image-text concatenation generates answer"""
        print("=" * 60)
        print(f"Question: {question}")
        print("=" * 60)

        # 1. Encode image and text
        print("\n1. Encode image...)
        vision_emb = self.encode_vision(image)

        print("2. Encode text...")
        text_emb = self.encode_text(question)

        # 3. Concatenate: [vision sequence] + [text sequence]
        print("\n3. Concatenate vision and text sequences...)
        print(f" Vision embedding shape: {vision_emb.shape}")
        print(f" Text embedding shape: {text_emb.shape}")
        combined = torch.cat([vision_emb, text_emb], dim=1)
        print(f" Shape after concatenation: {combined.shape}")

        # 4. LLM generates answer
        print("\n4. LLM generates answer...)

        return "This is a cute orange cat, it looks very happy."


# Demonstrate VLM Q&A
vlm = SimpleVLM()
answer = vlm.generate(
    image="A photo of an orange cat",
    question="What is in this picture? Please describe it."
)

print(f"\nAnswer: {answer}")


# ============================================
# Code example for actually using LLaVA
# ============================================
print("\n" + "=" * 60)
print("Code example for actually using LLaVA:")
print("=" * 60)
print("""
# LLaVA requires a lot of GPU memory, it is recommended to use the transformers or llava library

from transformers import AutoProcessor, AutoModelForVisionAndLanguageGeneration
import torch
from PIL import Image

# 1. Load model and processor
model_id = "llava-hf/llava-1.5-7b-hf"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForVisionAndLanguageGeneration.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

# 2. Prepare input
image = Image.open("cat.jpg")
prompt = "USER: <image>\\nWhat is in this picture? Please describe it. ASSISTANT:"

inputs = processor(images=image, text=prompt, return_tensors="pt").to("cuda")

# 3. Generate
with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=500)
answer = processor.decode(output[0], skip_special_tokens=True)

print(answer)
"""
)

Hands-On: Complete Project for Image Classification + Object Detection

Now let's put together the knowledge we've learned to build a complete hands-on project.

Example

# ============================================
# Hands-on Project: Complete Vision AI Pipeline
# Includes image classification, object detection, and result visualization
# ============================================

import torch
import torch.nn as nn
import torch.nn.functional as F
from torch.utils.data import Dataset, DataLoader
import numpy as np
from PIL import Image
import json
from pathlib import Path


class VisionPipeline:
    """
Complete Vision AI Pipeline

Features:
1. Image classification
2. Object detection
3. Result summary and visualization
    """


    def __init__(self, device: str = "cpu"):
        self.device = device
        self.classifier = None
        self.detector = None
        print(f"EXAMPLE Vision Pipeline initialized (device: {device})")

    def build_classifier(self, num_classes: int = 10):
        """Build image classifier"""
        print("\nBuilding image classifier...)
        # Use the previously defined SimpleCNN
        self.classifier = SimpleCNN(num_classes=num_classes).to(self.device)
        print("Image classifier built successfully!")
        return self

    def build_detector(self):
        """Build object detector"""
        print("\nBuilding object detector...)
        self.detector = SimpleYOLO()
        print("Detector built successfully!")
        return self

    def load_pretrained_weights(self, classifier_path: str = None):
        """Load pretrained weights"""
        if classifier_path and self.classifier:
            print(f"\nLoad classifier weights: {classifier_path}")
            # In practice: self.classifier.load_state_dict(torch.load(classifier_path))
        return self

    def classify_image(self, image: np.ndarray) -> tuple:
        """
Image Classification

Returns:
            (pred_class_id, pred_class_name, confidence)
        """

        if self.classifier is None:
            raise ValueError("Please build the classifier first: build_classifier()")

        # Preprocessing
        if isinstance(image, np.ndarray):
            if image.ndim == 2:  # Grayscale
                image = np.stack([image]*3, axis=-1)
            image = torch.from_numpy(image).permute(2, 0, 1).float() / 255.0
            image = F.interpolate(image.unsqueeze(0), size=(32, 32), mode='bilinear')
            image = image.to(self.device)

        # Inference
        self.classifier.eval()
        with torch.no_grad():
            logits = self.classifier(image)
            probs = F.softmax(logits, dim=-1)
            confidence, pred_idx = probs.max(dim=-1)

        # Return result
        class_names = [
            'airplane', 'automobile', 'bird', 'cat', 'deer',
            'dog', 'frog', 'horse', 'ship', 'truck'
        ]
        pred_class = class_names[pred_idx.item()]

        return pred_idx.item(), pred_class, confidence.item()

    def detect_objects(self, image: np.ndarray) -> list:
        """Object Detection"""
        if self.detector is None:
            raise ValueError("Please build the detector first: build_detector()")

        # Use simulated detection
        return self.detector.predict_dummy(image)

    def process_image(self, image: np.ndarray) -> dict:
        """
Process an entire image

Returns:
Dictionary containing all results
        """

        print("\n" + "=" * 60)
        print("Start processing image")
        print("=" * 60)

        results = {}

        # 1. Image classification
        if self.classifier is not None:
            print("\n[1/3] Running image classification...")
            class_id, class_name, conf = self.classify_image(image)
            results['classification'] = {
                'class_id': class_id,
                'class_name': class_name,
                'confidence': conf
            }
            print(f" → Classification result: {class_name} (Confidence: {conf:.2%})")

        # 2. Object detection
        if self.detector is not None:
            print("\n[2/3] Running object detection...")
            detections = self.detect_objects(image)
            results['detections'] = detections
            print(f" → Detected {len(detections)} objects:")
            for det in detections:
                print(f"      - {det['class_name']} ({det['score']:.2f})")

        # 3. Visualization
        print("\n[3/3] Generating visualization...")
        if self.detector is not None:
            vis_image = self.detector.visualize(image, detections)
            results['visualization'] = vis_image
            print(" → Visualization complete")

        print("\n" + "=" * 60)
        print("Image processing complete!")
        print("=" * 60)

        return results

    def save_results(self, results: dict, output_dir: str = "output"):
        """Save results to file"""
        output_path = Path(output_dir)
        output_path.mkdir(exist_ok=True)

        # Save JSON results
        json_result = {
            'classification': results.get('classification'),
            'detections': [
                {k: v for k, v in det.items() if k != 'bbox'}
                for det in results.get('detections', [])
            ]
        }
        with open(output_path / "results.json", "w") as f:
            json.dump(json_result, f, indent=2, ensure_ascii=False)

        # Save visualization image
        if 'visualization' in results:
            vis = Image.fromarray(results['visualization'])
            vis.save(output_path / "visualization.jpg")

        print(f"\nResults saved to: {output_path}/")


# Complete training example
class SimpleDataset(Dataset):
    """Simple simulated dataset"""

    def __init__(self, num_samples: int = 100):
        self.num_samples = num_samples
        # Generate fake data
        self.images = torch.randn(num_samples, 3, 32, 32)
        self.labels = torch.randint(0, 10, (num_samples,))

    def __len__(self):
        return self.num_samples

    def __getitem__(self, idx):
        return self.images[idx], self.labels[idx]


def train_classifier(pipeline: VisionPipeline, num_epochs: int = 5):
    """Complete workflow for training the classifier"""
    print("\n" + "=" * 60)
    print("Start training classifier")
    print("=" * 60)

    # 1. Prepare data
    print("\nPreparing dataset...")
    train_dataset = SimpleDataset(1000)
    val_dataset = SimpleDataset(200)

    train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True)
    val_loader = DataLoader(val_dataset, batch_size=32)

    print(f"Training set: {len(train_dataset)} samples")
    print(f"Validation set: {len(val_dataset)} samples")

    # 2. Optimizer and loss function
    optimizer = torch.optim.Adam(pipeline.classifier.parameters(), lr=1e-3)
    criterion = nn.CrossEntropyLoss()

    # 3. Training loop
    print("\nStart training...")
    for epoch in range(num_epochs):
        pipeline.classifier.train()
        total_loss = 0.0
        correct = 0
        total = 0

        for images, labels in train_loader:
            images, labels = images.to(pipeline.device), labels.to(pipeline.device)

            optimizer.zero_grad()
            outputs = pipeline.classifier(images)
            loss = criterion(outputs, labels)
            loss.backward()
            optimizer.step()

            total_loss += loss.item()
            _, predicted = outputs.max(1)
            total += labels.size(0)
            correct += predicted.eq(labels).sum().item()

        train_acc = correct / total
        avg_loss = total_loss / len(train_loader)

        # Validation
        pipeline.classifier.eval()
        correct = 0
        total = 0
        with torch.no_grad():
            for images, labels in val_loader:
                images, labels = images.to(pipeline.device), labels.to(pipeline.device)
                outputs = pipeline.classifier(images)
                _, predicted = outputs.max(1)
                total += labels.size(0)
                correct += predicted.eq(labels).sum().item()

        val_acc = correct / total

        print(f"Epoch {epoch+1}/{num_epochs} - "
              f"Loss: {avg_loss:.4f}, "
              f"Training accuracy: {train_acc:.2%}, "
              fValidation accuracy: {val_acc:.2%})

    print("\nTraining complete!)


# ============================================
Run the complete hands-on project.
# ============================================

print("=" * 60)
print(EXAMPLE Visual AI Complete Practical Project)
print("=" * 60)

# 1. Create Pipeline
pipeline = VisionPipeline(device="cuda" if torch.cuda.is_available() else "cpu")

# 2. Build the Model
pipeline.build_classifier(num_classes=10)
pipeline.build_detector()

3. Training the classifier
train_classifier(pipeline, num_epochs=3)

# 4. Processing Images
test_image = create_test_image()
results = pipeline.process_image(test_image)

5. Save the result
pipeline.save_results(results, output_dir="example_output")

print("\n" + "=" * 60)
print(Project run completed!)
print("=" * 60)
print("\nNext step learning suggestions:)
print(1. Try training on a real dataset (e.g., CIFAR-10).)
print(2. Fine-tuning using a pretrained model)
print(3. Try more complex model architectures.)
print(4. Deploy to practical applications.)
Other extensions