AI Multimodal

AI can not only talk, but also see images, hear sounds, and understand videos—master multimodal tools and technologies.

For a long time, AI was single-sensory.

  • Text AI could only process text, such as asking ChatGPT to write articles.

  • Image AI could only process pictures, such as asking Midjourney to draw images.

  • Audio AI could only process sound, such as asking Whisper to transcribe speech.

But the real world is multimodal—when we read books we see illustrations, when watching movies there are visuals and sound, and when communicating with others we observe expressions and tone.

Multimodal AI is AI that can simultaneously understand and generate multiple types of content, like giving AI eyes, ears, and a mouth, allowing it to interact with the world in a way closer to humans.

A practical example: you give AI a photo of a shopping receipt, it can recognize the text on it (OCR), understand what items were purchased, help you organize it into a table, and even read it aloud to you. This is a typical multimodal task.


What is Multimodal AI

First, let's clarify a few basic concepts.

Single-modal vs Multimodal

Modality refers to the form in which information is expressed.

Text, images, audio, and video are all different modalities.

TypeDescriptionTypical Products
Single-modal AICan only process one modalityEarly GPT-3 (text only), Stable Diffusion (images only)
Multimodal AICan process multiple modalities simultaneouslyGPT-4o、Claude 3、Gemini

The core breakthrough of multimodal AI is:It can convert information from different modalities into a common language for understanding。

For example, when seeing an image of a "cat", it can convert the image into a vector (a set of numbers); when seeing the word "cat", it can also convert it into a vector. These two vectors are close in mathematical space because they represent the same concept.

Core Challenges of Multimodal AI

Multimodality sounds simple, but in practice it is very difficult.

  • The first challenge is "alignment"—how to make the "cat" in an image and the "cat" in text appear as the same thing to the model?

  • The second challenge is "fusion"—when seeing an image and a piece of text at the same time, how do you combine their information?

  • The third challenge is "generation"—how to generate an image based on a text description, or generate a text description based on an image?

  • Fortunately, these problems are being gradually solved. After 2024, almost all mainstream large models are multimodal.


Image Understanding (Vision)

Image understanding is about making AI understand images—describing content, recognizing text, analyzing charts, and spotting details.

Mainstream large models now support image input.

There are usually two ways to use it:

  • One is to upload the image file directly, for example by clicking the image icon in the ChatGPT web version to upload.

  • Another way is to embed the image in the API request using Base64 encoding, suitable for programmatic calls.

ModelImage input methodsSupported image formats
GPTURL, Base64, direct uploadJPG、PNG、WEBP、GIF
ClaudeBase64, direct uploadJPG、PNG、WEBP、GIF
GeminiDirect upload, Google DriveJPG、PNG、WEBP

Common vision tasks include:

  • Image description — What is in this picture?

  • OCR — Extract the text from an image.

  • Chart analysis — What is this bar chart telling us?

  • Detail detection — Help me check this design draft for issues.

Hands-on: Calling the GPT-4o Vision API with Python

Let's look at a complete example — using the API to analyze an image.

Example

# ============================================
# File: example_vision_demo.py
# Function: Analyze images using the GPT-4o Vision API
# ============================================

import base64
import requests
import os

# Configure API key (please replace with your real key)
# Get it from: https://platform.openai.com/api-keys
OPENAI_API_KEY = "sk-your-api-key-here"
OPENAI_API_URL = "https://api.openai.com/v1/chat/completions"


def encode_image_to_base64(image_path: str) -> str:
    """
Encode the local image file as a Base64 string
This is the transfer format required by the API
    """

    with open(image_path, "rb") as image_file:
        # Read the binary file content and encode it with Base64
        base64_data = base64.b64encode(image_file.read()).decode("utf-8")
        return base64_data


def analyze_image_with_gpt4o(
    image_path: str,
    prompt: str = "Please describe the content of this image in detail.",
    model: str = "gpt-4o"
) -> str:
    """
Use the GPT-4o Vision API to analyze the image

Parameter description:
image_path: Local image file path (e.g., "example_test.jpg")
prompt: The question you want to ask or the task you want the AI to perform
model: The model name to use (default gpt-4o)

Returns:
The AI's analysis result text
    """

    # Step 1: Encode the image as Base64
    base64_image = encode_image_to_base64(image_path)

    # Step 2: Build the request headers
    headers = {
        "Content-Type": "application/json",
        "Authorization": f"Bearer {OPENAI_API_KEY}"
    }

    # Step 3: Build the request body
    # Note: The image content is passed via the image_url field, in the format "data:image/jpeg;base64,..."
    payload = {
        "model": model,
        "messages": [
            {
                "role": "user",
                "content": [
                    # The first part is the text prompt
                    {"type": "text", "text": prompt},
                    # The second part is the image
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/jpeg;base64,{base64_image}"
                        }
                    }
                ]
            }
        ],
        "max_tokens": 1000  # Limit output length to avoid high costs
    }

    # Step 4: Send the API request
    print(f"Analyzing image: {image_path} ...")
    response = requests.post(OPENAI_API_URL, headers=headers, json=payload)

    # Step 5: Handle the response
    if response.status_code == 200:
        result = response.json()
        # Extract the AI's answer
        answer = result["choices"][0]["message"]["content"]
        return answer
    else:
        # Print detailed error information when an error occurs
        error_msg = f"Request failed: {response.status_code} - {response.text}"
        print(error_msg)
        return error_msg


def extract_text_from_image(image_path: str) -> str:
    """
Specially for OCR: Extract text from images
    """

    prompt = """
Please extract all text content from this image.
Requirements:
1. Accurately reproduce the text, do not omit anything
2. Keep the original paragraphs and formatting
3. If there are tables, output them in table form
4. If there is no text in the image, please state so
This is the OCR test from the EXAMPLE tutorial.
    """

    return analyze_image_with_gpt4o(image_path, prompt)


def analyze_chart(image_path: str) -> str:
    """
Specially for chart analysis
    """

    prompt = """
Please analyze this chart.
Please answer the following questions:
1. What type of chart is this (bar chart, line chart, pie chart, etc.)?
2. What are the chart title and the meanings of the axes?
3. What is the main message conveyed by the chart?
4. Are there any notable trends or anomalies?
This is the chart analysis test from the EXAMPLE tutorial.
    """

    return analyze_image_with_gpt4o(image_path, prompt)


# ============================================
# Main program: Demonstrating different vision analysis tasks
# ============================================

if __name__ == "__main__":
    # Assume we have a test image (you need to prepare a real image)
    # It can be: product photo, screenshot, document photo, chart, etc.
    test_image = "example_test_image.jpg"

    # Check if the image exists
    if not os.path.exists(test_image):
        print(f"Tip: Please prepare an image, name it {test_image}, and place it in the current directory")
        print("Or modify the image path in the code")
    else:
        # Task 1: General image description
        print("\n" + "="*50)
        print("Task 1: Image description")
        print("="*50)
        result1 = analyze_image_with_gpt4o(
            test_image,
            "Please describe this image in detail, including the subject, background, colors, composition, etc."
        )
        print(result1)

        # Task 2: OCR text extraction
        print("\n" + "="*50)
        print("Task 2: OCR Text Extraction")
        print("="*50)
        result2 = extract_text_from_image(test_image)
        print(result2)

        # Task 3: Chart Analysis (if the image is a chart)
        print("\n" + "="*50)
        print("Task 3: Chart Analysis")
        print("="*50)
        result3 = analyze_chart(test_image)
        print(result3)

        print("\n" + "="*50)
        print("EXAMPLE tutorial demo complete!")
        print("="*50)

Before running this program, you need to:

  • 1. Install dependencies:pip install requests

  • 2. Prepare a test image and name itexample_test_image.jpg

  • 3. Replace the API key with your own

Cost note: GPT-4o Vision is billed by image size. A 1024x1024 image costs about $0.00765. Smaller images are cheaper. For debugging, use low-resolution images; for production, use high resolution.


Text-to-Image

Text-to-Image lets AI generate images from text descriptions. This is another very mature multimodal application area.

Introduction to Diffusion Model Principles

Today's text-to-image tools basically use diffusion model technology.

A simple way to understand how it works:

  • Training phase: Show the model a clear image, then gradually add noise to it until it becomes a completely noisy image. The model learns how to turn the noisy image back into a clear one step by step.

  • Generation phase: Give the model a random noise image, and let it denoise step by step according to the text prompt, finally generating a clear image that matches the description.

This process usually takes 20-50 steps, so generating an image takes a few seconds.

Advanced Midjourney Usage Tips

Midjourney is currently one of the most popular text-to-image tools, used via Discord.

A high-quality prompt usually includes these elements:

  • Subject description — "a cat wearing a hat"

  • Style — "Studio Ghibli style", "oil painting", "3D render"

  • Lighting — "natural light", "cinematic lighting", "cyberpunk neon lights"

  • Composition — "close-up", "wide angle", "top-down view"

  • Quality — "8K", "ultra-high detail", "photographic quality"

ParameterDescriptionExample
--arSet aspect ratio--ar 16:9 (video), --ar 3:4 (mobile)
--vSpecify model version--v 6.0 (latest version)
--sStylization level (0-1000)--s 750 (strong artistic feel)
--qRender quality (0.25-2)--q 2 (higher quality, slower)
--iwReference image weight--iw 1.5 (more like the reference image)

Local Deployment of Stable Diffusion

If you want full control, you can run Stable Diffusion locally:https://github.com/compvis/stable-diffusion。

The most commonly used tool is Automatic1111's WebUI.

The steps are roughly:

  • 1. Install Python 3.10

  • 2. Download Stable Diffusion WebUI

  • 3. Download the model file (.safetensors format)

  • 4. Run webui.bat (Windows) or webui.sh (Mac/Linux)

Advantages of local deployment: generation is free, no content restrictions, can use LoRA to fine-tune styles, and can use ControlNet for precise composition control.

Disadvantages: you need a GPU with at least 6GB of VRAM, and setup is somewhat complicated.

Hands-on: Generating Images with the DALL·E 3 API

If you want a simple API way to generate images, DALL·E 3 is a good choice.

Example

# ============================================
# File: example_dalle3_demo.py
# Function: Generate images using the DALL·E 3 API
# ============================================

import requests
import os
from datetime import datetime

# Configure the API key (please replace with your real key)
OPENAI_API_KEY = "sk-your-api-key-here"
DALLE3_API_URL = "https://api.openai.com/v1/images/generations"


def generate_image_with_dalle3(
    prompt: str,
    size: str = "1024x1024",
    quality: str = "standard",
    style: str = "vivid",
    save_dir: str = "example_generated_images"
) -> dict:
    """
Generating Images with the DALL·E 3 API

Parameter descriptions:
prompt: text description, the more detailed the better
size: image size
Options: 1024x1024, 1024x1792 (portrait), 1792x1024 (landscape)
quality: image quality
Options: standard (standard), hd (high definition, more expensive)
style: style
Optional: vivid (vivid, more artistic), natural (natural, more realistic)
save_dir: image save directory

Returns:
A dictionary containing the image URL and save path
    """

    # Step 1: Prepare the request headers
    headers = {
        "Content-Type": "application/json",
        "Authorization": f"Bearer {OPENAI_API_KEY}"
    }

    # Step 2: Prepare the request body
    payload = {
        "model": "dall-e-3",
        "prompt": prompt,
        "n": 1,  # DALL·E 3 can only generate 1 image at a time
        "size": size,
        "quality": quality,
        "style": style
    }

    # Step 3: Send the request
    print(f"Generating image, prompt: {prompt[:50]}...")
    response = requests.post(DALLE3_API_URL, headers=headers, json=payload)

    # Step 4: Handle the response
    if response.status_code == 200:
        result = response.json()
        image_url = result["data"][0]["url"]
        revised_prompt = result["data"][0].get("revised_prompt", prompt)

        print(f"DALL·E 3 auto-optimized prompt: {revised_prompt}")

        # Step 5: Download and save the image
        if not os.path.exists(save_dir):
            os.makedirs(save_dir)

        # Generate a filename using a timestamp
        timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
        filename = f"example_dalle3_{timestamp}.png"
        filepath = os.path.join(save_dir, filename)

        # Download the image
        image_response = requests.get(image_url)
        if image_response.status_code == 200:
            with open(filepath, "wb") as f:
                f.write(image_response.content)
            print(f"Image saved to: {filepath}")

        return {
            "image_url": image_url,
            "filepath": filepath,
            "revised_prompt": revised_prompt
        }
    else:
        error_msg = f"Request failed: {response.status_code} - {response.text}"
        print(error_msg)
        return {"error": error_msg}


def generate_example_logo_demo():
    """Demo: Generate a EXAMPLE-style Logo"""
    prompt = """
    A cute and friendly robot mascot for EXAMPLE tutorial website.
    Blue and green color scheme, modern flat design, white background.
    The robot looks helpful and intelligent, with big friendly eyes.
    Simple, clean, suitable for a tech education brand.
    """

    # Chinese prompts work too; DALL·E 3 supports multiple languages
    prompt = """
Design a cute and friendly robot mascot for the EXAMPLE tutorial website.
Blue-green color scheme, modern flat design, white background.
The robot looks helpful and smart, with big friendly eyes.
Simple and clean, suitable for a tech education brand.
    """

    return generate_image_with_dalle3(
        prompt=prompt,
        size="1024x1024",
        quality="hd",
        style="vivid"
    )


def generate_book_cover_demo():
    """Demo: Generate a book cover"""
    prompt = """
    Book cover design for "AI Tutorial for Beginners".
    Modern minimalist style, deep blue and purple gradient background.
    The title text is elegant and readable.
    Subtle tech elements: floating geometric shapes, faint circuit patterns.
    Professional, trustworthy, suitable for a programming book.
    """

    return generate_image_with_dalle3(
        prompt=prompt,
        size="1024x1792",  # Portrait orientation
        quality="hd",
        style="natural"
    )


def generate_banner_demo():
    """Demo: Generate a banner ad"""
    prompt = """
    Web banner design for a coding tutorial website (EXAMPLE).
    16:9 wide format. Modern gradient background from teal to blue.
    Clean layout with space for text.
    Abstract tech elements: floating code snippets, subtle grid lines.
    Professional and inviting.
    """

    return generate_image_with_dalle3(
        prompt=prompt,
        size="1792x1024",  # Landscape orientation
        quality="standard",
        style="vivid"
    )


# ============================================
# Main program: demonstrate different generation tasks
# ============================================

if __name__ == "__main__":
    print("="*60)
    print("EXAMPLE DALL·E 3 Image Generation Demo")
    print("="*60)

    # Task 1: Generate Logo
    print("\n[Task 1] Generating Logo ...")
    result1 = generate_example_logo_demo()

    # Task 2: Generate book cover
    print("\n[Task 2] Generating book cover ...")
    result2 = generate_book_cover_demo()

    # Task 3: Generate banner
    print("\n[Task 3] Generating banner ...")
    result3 = generate_banner_demo()

    print("\n" + "="*60)
    print("All tasks completed! Check the example_generated_images directory")
    print("="*60)

DALL·E 3's strengths: strong understanding, doesn't break hands and faces, and generates text well.

But it also has limitations: it can only generate 1 image at a time, cannot precisely control style, and cannot be fine-tuned like Stable Diffusion.

Copyright notice: The current consensus is that for AI-generated images, the prompt writer has usage rights but does not own the copyright (because copyright offices require human authorship). For commercial use, it is recommended to confirm the tool's terms of use.


Video AI

Video AI is divided into two categories: video generation and video understanding.

Compared to images, video is technically more difficult—because video is a combination of "time + space," requiring every frame to be coherent and movements to look natural.

Comparison of AI Video Generation Tools

2024 was the first year of AI video, with multiple companies launching video generation products.

ToolDeveloperFeaturesTypical duration
SoraOpenAIIndustry-leading image detail and spatial scene understanding, realistic physical logic, not yet fully open for public betaUp to 60 seconds
Runway Gen-4RunwayAll-in-one creation platform, integrating text/image-to-video, AI post-editing, motion brush, and VFX effects; the top choice for overseas professional creators10–20 seconds
Kling (Keling AI)KuaishouExcellent adaptation to Chinese prompts, smooth character movements, stable characters across shots, high cost-effectiveness in China, supports multi-segment narrative stitching5–30 seconds, up to 2 minutes per segment
PikaPika LabsExtremely expressive in anime and cinematic illustration styles, consistent character design, supports video style redrawing and frame interpolation5–10 seconds
VeoGoogleCinematic lighting and camera movement, stable long-take narrative, natural physical rendering, not yet publicly releasedOver 60 seconds
Jimeng AI (Seedance)ByteDanceDeep integration with the CapCut/Douyin ecosystem, precise Chinese semantic understanding, audio-visual lip sync, supports automatic storyboarding from scripts5–15 seconds, multiple segments stitched into a complete film
Tongyi Wanxiang VideoAlibaba DAMO AcademyOpen-source base model, stable multi-subject interaction, supports text generation in images, can be deployed locally for inference.2–15 seconds
Hailuo AI (Hailuo)MiniMax (Xiyu)Lightweight ultra-fast generation; suitable for converting posters into dynamic short videos and social media flash content, with outstanding anime-style rendering.4–10 seconds
Luma Dream MachineLuma AITop-notch physics dynamics simulation; realistic effects for water, fire, cloth, and explosions, with cinematic camera movement.5–12 seconds
ZhiyingTencentFocuses on AI editing and digital-human videos, with massive short-video templates, one-stop subtitles, dubbing, and matting, and low commercial barriers.No limit on video length; AI-generated clips are 5–15 seconds.
Haiyi AIHaiyi TechnologySupports 4K/60fps high-frame-rate output, multi-image linked video generation, comes with a complete production workstation, and provides ample time-limited free quota.Single segment up to 30 seconds.

AI video is still in its early stages:

  • Advantages: quick creativity, low cost, and the ability to achieve effects that are very difficult with live-action shooting.

  • Disadvantages: limited duration, details easily break down, and physics logic is sometimes incorrect (e.g., objects passing through each other, wrong number of fingers).

  • Suitable scenarios: concept demonstrations, short-video material, e-commerce ads, game animations.

  • Unsuitable scenarios: long films requiring precise control, serious news, legal evidence.

Video Understanding and Summarization

Besides generating videos, AI can also understand videos.

Common tasks in video understanding:

  • Video summarization — "What is this video about? Summarize in 3 sentences."

  • Content retrieval — "Find the clip in the video where someone is riding a bicycle."

  • Question answering — "What did the third person in the video say?"

The typical technical approach is: extract frames from the video (e.g., 1 frame per second), then use a vision model to look at each frame one by one, and finally integrate the information.

Future Trends

The development direction of video AI is clear:

  • Longer — from 10 seconds to 1 minute, and then to 10 minutes.

  • More controllable — able to precisely control every shot using storyboards.

  • More consistent — the same character stays consistent across different shots.

  • More practical — from looking cool to solving real-world problems.

  • What to expect: in a few years, everyone may be able to create film-like video works with AI.


Audio and Speech AI

Audio AI mainly includes: speech-to-text (ASR), text-to-speech (TTS), and music generation.

This is a fairly mature field.

Speech-to-Text (Whisper API)

Whisper is OpenAI's open-source speech recognition model. It performs very well and supports 99 languages.

Example

# ============================================
# File: example_whisper_demo.py
# Function: Use Whisper API to transcribe speech
# ============================================

import requests
import os

# Configure API key
OPENAI_API_KEY = "sk-your-api-key-here"
WHISPER_API_URL = "https://api.openai.com/v1/audio/transcriptions"


def transcribe_audio_with_whisper(
    audio_path: str,
    language: str = "zh",  # zh=Chinese, en=English, leave empty for auto-detect
    prompt: str = "",     # Optional: provide some terminology hints to improve accuracy
    temperature: float = 0.0
) -> dict:
    """
Use Whisper API to transcribe audio files

Parameter description:
audio_path: audio file path (supports mp3, wav, m4a, etc.)
language: language code (optional, auto-detected if not specified)
Common codes: zh (Chinese), en (English), ja (Japanese)
prompt: prompt text (optional; write possible technical terms or names here)
temperature: sampling temperature (0.0 is most deterministic, 1.0 is more diverse)

Return:
A dictionary containing the transcribed text
    """

    # Step 1: Check if the file exists
    if not os.path.exists(audio_path):
        return {"error": f"File not found: {audio_path}"}

    # Step 2: Prepare the request headers
    headers = {
        "Authorization": f"Bearer {OPENAI_API_KEY}"
    }

    # Step 3: Prepare the request body (note: audio must use multipart/form-data format)
    files = {
        "file": open(audio_path, "rb")
    }
    data = {
        "model": "whisper-1",
        "temperature": temperature
    }
    if language:
        data["language"] = language
    if prompt:
        data["prompt"] = prompt

    # Step 4: Send the request
    print(f"Transcribing audio: {audio_path} ...")
    response = requests.post(WHISPER_API_URL, headers=headers, files=files, data=data)

    # Step 5: Close the file
    files["file"].close()

    # Step 6: Process the response
    if response.status_code == 200:
        result = response.json()
        text = result["text"]
        print(f"Transcription complete, a total of {len(text)} characters.")
        return {
            "text": text,
            "language": language
        }
    else:
        error_msg = f"Request failed: {response.status_code} - {response.text}"
        print(error_msg)
        return {"error": error_msg}


def transcribe_with_timestamps(audio_path: str) -> dict:
    """
Timestamped transcription (using verbose_json format)
    """

    headers = {
        "Authorization": f"Bearer {OPENAI_API_KEY}"
    }
    files = {"file": open(audio_path, "rb")}
    data = {
        "model": "whisper-1",
        "response_format": "verbose_json",  # Returns detailed information, including timestamps
        "timestamp_granularities": ["segment", "word"]  # Supports segment-level and word-level timestamps
    }

    response = requests.post(WHISPER_API_URL, headers=headers, files=files, data=data)
    files["file"].close()

    if response.status_code == 200:
        result = response.json()
        return result
    else:
        return {"error": response.text}


def translate_audio_to_english(audio_path: str) -> dict:
    """
Translate audio from any language into English text
    """

    headers = {
        "Authorization": f"Bearer {OPENAI_API_KEY}"
    }
    files = {"file": open(audio_path, "rb")}
    data = {
        "model": "whisper-1",
        "task": "translate"  # Specifically specify: translation task
    }

    response = requests.post("https://api.openai.com/v1/audio/translations",
                            headers=headers, files=files, data=data)
    files["file"].close()

    if response.status_code == 200:
        return response.json()
    else:
        return {"error": response.text}


def example_meeting_minutes_demo(audio_path: str) -> str:
    """
Hands-on: generating meeting minutes
Step 1: Transcribe the recording with Whisper
Step 2: Organize into minutes with GPT-4o
    """

    # First transcribe
    print("Step 1: Transcribing the meeting recording...")
    transcribe_result = transcribe_audio_with_whisper(
        audio_path,
        language="zh",
        prompt="This is a meeting of the EXAMPLE technical team, discussing product development, progress, technical solutions, etc."
    )

    if "error" in transcribe_result:
        return transcribe_result["error"]

    transcript = transcribe_result["text"]

    # Then organize (simplified here; in a real project you can call the GPT-4o API)
    print("Step 2: Organizing the meeting minutes...")
    # For demonstration purposes, return the transcript text directly
    # In a real project, you can send the transcript to GPT-4o and have it organize it
    return transcript


# ============================================
# Main program
# ============================================

if __name__ == "__main__":
    # Prepare a test audio file
    test_audio = "example_test_audio.mp3"

    if not os.path.exists(test_audio):
        print(f"Hint: Prepare an audio file and name it {test_audio}")
        print("Or record a voice memo on your phone and save it as MP3")
    else:
        # Task 1: Simple transcription
        print("\n" + "="*50)
        print("Task 1: Simple transcription")
        print("="*50)
        result1 = transcribe_audio_with_whisper(test_audio, language="zh")
        if "text" in result1:
            print("Transcription result:")
            print(result1["text"])

        # Task 2: Timestamped transcription
        print("\n" + "="*50)
        print("Task 2: Timestamped transcription")
        print("="*50)
        result2 = transcribe_with_timestamps(test_audio)
        if "segments" in result2:
            print("Segment information:")
            for seg in result2["segments"][:3]:  # Only print the first 3 segments
                start = round(seg["start"], 2)
                end = round(seg["end"], 2)
                text = seg["text"]
                print(f"[{start}-{end}] {text}")

        print("\n" + "="*50)
        print("EXAMPLE Whisper demo completed!")
        print("="*50)

Whisper's advantages are: good multilingual support, not picky about accents, and automatic punctuation.

Typical application scenarios: meeting recording transcription, video subtitle generation, podcast transcripts, voice diaries.

Text-to-Speech (TTS)

Text-to-speech lets AI "read" articles.

OpenAI also provides a TTS API supporting multiple voice styles.

Other TTS tools: ElevenLabs (most human-like), Azure TTS (commercial-grade), Edge TTS (free), Coqui (open-source).

Music Generation (Suno, Udio)

Music generation has also achieved breakthroughs.

Suno and Udio are two representative products—you input lyrics and style descriptions, and they can generate a complete song with vocals and accompaniment.

Capability boundaries:

  • Can do: generate background music, write song demos, quickly try styles.

  • Limited in: fully replicating a specific song, precisely controlling notes, professional-grade post-production.

  • Copyright: terms vary by tool; please confirm before commercial use.


Introduction to Multimodal Model Architectures

Understanding a bit of the theory helps you use the tools better.

CLIP: The Foundation of Image-Text Alignment

CLIP (Contrastive Language-Image Pre-training) is a model published by OpenAI in 2021.

Its idea is simple: look at images and text simultaneously, and learn to associate them.

During training, give the model a batch ofimage + descriptionpairs, and let it learn:

  • 1. Convert images into vectors

  • 2. Convert text into vectors

  • 3. Make paired image-text vectors as close as possible, and unpaired ones as far apart as possible

CLIP's significance lies in:It proved that "cross-modal understanding" is feasible。

Later models like Stable Diffusion and GPT-4o were all inspired by CLIP.

LLaVA: Open-Source Multimodal LLM

LLaVA (Large Language and Vision Assistant) is an open-source multimodal model.

Its architecture is very clear:

  • A vision encoder (e.g., ViT) — responsible for seeing images

  • A large language model (e.g., LLaMA) — responsible for speaking

  • A projection layer — converts image vectors into a format the language model can understand

Training is done in two steps:

  • Step 1: Pretraining alignment — make the vision encoder and language model speak the same language

  • Step 2: Instruction fine-tuning - train with image-question-answer data to teach it to answer according to instructions

LLaVA is a good starting point for learning multimodal technology - open-source code, clear architecture, good results.


Hands-on: Building an Image Analysis Tool

Let's integrate the previous knowledge and build a practical image analysis tool.

Example

# ============================================
# File: example_image_analyzer.py
# Function: a complete multimodal image analysis tool
# ============================================

import base64
import requests
import os
import json
from datetime import datetime
from typing import Optional, Dict, Any


class ExampleImageAnalyzer:
    """
EXAMPLE Image Analyzer - a packaged multimodal tool class
    """


    def __init__(self, api_key: str, base_url: str = "https://api.openai.com/v1"):
        """
Initialize analyzer

Parameters:
api_key: OpenAI API key
base_url: API address (configurable proxy)
        """

        self.api_key = api_key
        self.base_url = base_url
        self.headers = {
            "Content-Type": "application/json",
            "Authorization": f"Bearer {api_key}"
        }

    def _encode_image(self, image_path: str) -> str:
        """Encode the image as Base64"""
        with open(image_path, "rb") as f:
            return base64.b64encode(f.read()).decode("utf-8")

    def _call_vision_api(
        self,
        image_path: str,
        prompt: str,
        model: str = "gpt-4o",
        max_tokens: int = 2000
    ) -> Optional[str]:
        """Internal method: call the vision API"""
        base64_image = self._encode_image(image_path)

        payload = {
            "model": model,
            "messages": [
                {
                    "role": "user",
                    "content": [
                        {"type": "text", "text": prompt},
                        {
                            "type": "image_url",
                            "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}
                        }
                    ]
                }
            ],
            "max_tokens": max_tokens
        }

        url = f"{self.base_url}/chat/completions"
        response = requests.post(url, headers=self.headers, json=payload)

        if response.status_code == 200:
            return response.json()["choices"][0]["message"]["content"]
        else:
            print(f"API error: {response.status_code} - {response.text}")
            return None

    def describe(self, image_path: str) -> Optional[str]:
        """Describe the image content in detail"""
        prompt = """
Please describe this image in detail, including:
1. Main content (what people/objects are there)
2. Background environment
3. Colors and lighting
4. Overall atmosphere
Answer in Chinese, with a clear structure.
        """

        return self._call_vision_api(image_path, prompt)

    def extract_text(self, image_path: str) -> Optional[str]:
        """Extract all text from the image (OCR)"""
        prompt = """
Please extract all text content from this image.
Requirements:
1. Reproduce accurately without omission
2. Preserve the original paragraph structure
3. If it is a table, output it in Markdown table format
4. If there is no text, explicitly state "No recognizable text in the image"
        """

        return self._call_vision_api(image_path, prompt)

    def analyze_ui(self, image_path: str) -> Optional[str]:
        """Analyze UI design mockup (screenshot)"""
        prompt = """
Please analyze this UI design mockup:
1. What type of interface is this (APP, web page, mini-program, etc.)?
2. What are the main functional modules of the interface?
3. What is the design style (color scheme, layout, typography)?
4. What could be improved?
Answer in Chinese, listed in points.
        """

        return self._call_vision_api(image_path, prompt)

    def analyze_product(self, image_path: str) -> Optional[str]:
        """Analyze product images and generate e-commerce copy"""
        prompt = """
Please analyze this product image and generate e-commerce copy.
The copy should include:
1. Product name and type
2. Appearance features (color, material, design)
3. An attractive short recommendation (30-50 characters)
Answer in Chinese, in a friendly and positive tone.
        """

        return self._call_vision_api(image_path, prompt)

    def classify(self, image_path: str, categories: list = None) -> Optional[str]:
        """Classify the image"""
        if categories is None:
            categories = [
                "Portrait photo", "Landscape photo", "Food photo", "Product photo",
                "Design mockup/screenshot", "Document/table", "Chart/data visualization", "Other"
            ]
        cat_list = "、".join(categories)
        prompt = f"""
Please determine which category this image belongs to.
Optional categories: {cat_list}
Only return the most matching category name, no other text.
        """

        return self._call_vision_api(image_path, prompt)

    def full_analysis(self, image_path: str, output_dir: str = "example_analysis") -> Dict[str, Any]:
        """
Full analysis: generate a detailed analysis report

Return a dict containing all analysis results
        """

        print(f"Start analyzing image: {image_path}")

        # Classify first
        print("[1/5] Classifying...")
        category = self.classify(image_path) or "Unknown"

        # Description
        print("[2/5] Describing...")
        description = self.describe(image_path) or ""

        # Extract text
        print("[3/5] OCR in progress...")
        text = self.extract_text(image_path) or ""

        # Perform specialized analysis based on category
        print("[4/5] Specialized analysis in progress...")
        special_analysis = ""
        if "design" in category or "screenshot" in category:
            special_analysis = self.analyze_ui(image_path) or ""
        elif "product" in category or "merchandise" in category:
            special_analysis = self.analyze_product(image_path) or ""

        # Save report
        print("[5/5] Generating report...")
        if not os.path.exists(output_dir):
            os.makedirs(output_dir)

        timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
        report = {
            "image_file": os.path.basename(image_path),
            "analysis_time": datetime.now().isoformat(),
            "category": category,
            "description": description,
            "extracted_text": text,
            "special_analysis": special_analysis
        }

        # Save JSON
        json_path = os.path.join(output_dir, f"report_{timestamp}.json")
        with open(json_path, "w", encoding="utf-8") as f:
            json.dump(report, f, ensure_ascii=False, indent=2)

        # Save readable text
        txt_path = os.path.join(output_dir, f"report_{timestamp}.txt")
        with open(txt_path, "w", encoding="utf-8") as f:
            f.write("="*60 + "\n")
            f.write("EXAMPLE Image Analysis Report\n")
            f.write("="*60 + "\n\n")
            f.write(f"Image: {report['image_file']}"\n")
            f.write(f"Time: {report['analysis_time']}"\n")
            f.write(f"Category: {report['category']}"\n\n")
            f.write("-"*60 + "\n")
            f.write("Image description:\n")
            f.write("-"*60 + "\n")
            f.write(report["description"] + "\n\n")
            f.write("-"*60 + "\n")
            f.write("Extracted text:\n")
            f.write("-"*60 + "\n")
            f.write(report["extracted_text"] + "\n\n")
            if report["special_analysis"]:
                f.write("-"*60 + "\n")
                f.write("Special analysis:\n")
                f.write("-"*60 + "\n")
                f.write(report["special_analysis"] + "\n")

        print(f"Analysis complete! Report saved to: {output_dir}")
        return report


# ============================================
# Usage example
# ============================================

def main():
    # Configure your API Key
    api_key = "sk-your-api-key-here"

    # Create an analyzer
    analyzer = ExampleImageAnalyzer(api_key)

    # Path of the image to analyze
    image_path = "example_test_image.jpg"

    if not os.path.exists(image_path):
        print(f"Please prepare the image: {image_path}")
        return

    # Method 1: Use a single function individually
    print("\n"Method 1: Use OCR alone")
    print("-"*40)
    text = analyzer.extract_text(image_path)
    if text:
        print("Extracted text:")
        print(text)

    # Method 2: Complete analysis
    print("\n"Method 2: Complete analysis")
    print("-"*40)
    report = analyzer.full_analysis(image_path)
    print(f"Classification result: {report['category']}")
    print(f"Report saved")


if __name__ == "__main__":
    main()

This tool encapsulates common image analysis functions and can be used directly in projects.

You can extend it as needed:

  • Add more analysis scenarios (e.g., medical imaging, industrial inspection, educational question analysis)

  • Add a simple web interface (using Flask or Gradio)

  • Support batch processing of images in folders

  • Store the results in a database

Other extensions