AI API Development

It is convenient to use AI chat tools in a web interface, but to truly integrate AI capabilities into your own products, workflows, and automation scripts, you need an API.

An API (Application Programming Interface) is a bridge for communication between applications. Through an API, your code can directly send requests to an AI model and receive replies, without needing to manually open a webpage and copy-paste.

Imagine these scenarios:

  • Your e-commerce website automatically generates descriptive copy for each product.

  • Your note-taking app automatically summarizes long text entered by users.

  • Your customer service system automatically categorizes and replies to user inquiries.

  • All these features can be implemented through AI APIs.

The goal of this module is to take you from "only being able to chat with AI" to "being able to make AI work in your code."


API Basic Concepts

Before writing code, first understand a few core concepts.

What is an API

An API is a set of agreed-upon communication protocols. We send requests in a specific format, and the other party returns results in a specific format.

Use ordering food as an analogy:

  • You walk into a restaurant, get the menu (API documentation), and know what you can order and how to order (request format).

  • You say to the waiter: "One Kung Pao Chicken, less spicy" (sending an API request).

  • The kitchen prepares the dish and brings it to you (returning the API response).

In this process, you don't need to enter the kitchen, and you don't need to know how the dish is made—this is the value of an API:Encapsulate complex details, expose only a simple interface。

REST API Fundamentals

REST (Representational State Transfer) is currently the most popular API design style.

Simply put, it uses the HTTP protocol for communication:

HTTP MethodPurposeExample in AI APIs
POSTCreate a resourceSend chat requests, generate images
GETRetrieve a resourceQuery model lists, get billing information
DELETEDelete a resourceDelete conversation history

The vast majority of AI APIs use POST requests, because we need to "send content to AI for processing."

JSON Format Basics

JSON (JavaScript Object Notation) is the most commonly used data format in API communication.

JSON's syntax is very simple; it's a combination of key-value pairs.

Examples

{
    "model": "gpt-4",
    "messages": [
        {
            "role": "user",
            "content": "Hello, please introduce EXAMPLE"
        }
    ],
    "temperature": 0.7,
    "max_tokens": 1000
}

This is a typical AI API request body.

Several core rules:

  • For objects, use{ }to wrap, for arrays use[ ]to wrap.

  • Strings use double quotes", cannot use single quotes.

  • Key-value pairs use a colon:to separate, keys must be strings.

  • Multiple key-value pairs use a comma,to separate.

JSON looks a lot like Python dictionaries, but there are syntactic differences; pay attention when writing code.


Development Environment Setup

To do a good job, one must first sharpen one's tools.

Python Installation and Configuration

Python is the most commonly used language for calling AI APIs, with a mature ecosystem and rich libraries.

First, check whether Python is installed on your computer:

# 检查 Python 版本(Windows、macOS、Linux 通用)
python --version
# 或者
python3 --version

It is recommended to use Python 3.9 or higher.

If not installed, go topython.orgDownload the latest stable version.

pip Package Management

pip is Python's package manager, used to install third-party libraries.

# 检查 pip 版本
pip --version
# 或者
pip3 --version

# 升级 pip 到最新版本
pip install --upgrade pip

# 配置国内镜像源(可选,提升下载速度)
pip config set global.index-url https://pypi.tuna.tsinghua.edu.cn/simple

Code Editor Selection

For writing AI API code, we recommend one of the following editors:

EditorFeaturesTarget Users
VS CodeFree, rich plugin ecosystem, lightweightMost developers
CursorBuilt-in AI assistance, can directly generate codeFor those who want AI to help write code
PyCharmMost feature-rich, best Python supportProfessional Python developers

The example code in this tutorial can run in any editor.


OpenAI API Basics

The OpenAI API is the most classic and well-established AI API, and many other vendors are compatible with its format.

Registration and Getting API Key

The API Key is your pass, proving your identity and recording your usage. Steps to get one:

  • 1. Visitplatform.openai.comRegister an account

  • 2. Go to the API Keys page

  • 3. Click "Create new secret key"

  • 4. Copy and save this Key (it is only shown once!)

Important:Never commit your API Key to a public code repository. Once leaked, others may use your quota and incur charges.

First API Call

First, install the official OpenAI SDK:

pip install openai

Now write your first API call program:

Examples

# File path: first_api_call.py
# First OpenAI API call example

from openai import OpenAI

# Initialize the client
# Note: It is recommended to read the API Key from environment variables; do not hardcode it in code
# For demonstration purposes, it is written directly here. In real projects, use environment variables.
client = OpenAI(
    api_key="your-api-key-here"  # Replace with your API Key
)

# Send a chat request
response = client.chat.completions.create(
    model="gpt-3.5-turbo",  # Select the model
    messages=[
        {
            "role": "user",  # Role: user represents the user
            "content": "Please introduce EXAMPLE's Rookie Tutorial in one sentence."  # User input content
        }
    ]
)

# Print the full response (see what was returned)
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")

# Extract the AI's reply
ai_reply = response.choices[0].message.content
print("AI reply:")
print(ai_reply)

If you don't have an OpenAI api_key, you can also use domestic alternatives. Many domestic models are compatible with OpenAI, such as DeepSeek. You can go tohttps://platform.deepseek.com/api_keysapply for an api_key and then replace yours with the following codeapi_keyandmodel, DeepSeek supports the models deepseek-v4-flash and deepseek-v4-pro (thinking mode), and you can use the API to call large models:

Examples

from openai import OpenAI

# Initialize the client
# Note: It is recommended to read the API Key from an environment variable, do not hardcode it in the code
# For demonstration purposes, it is written directly here. In actual projects, please use environment variables
client = OpenAI(
    api_key="sk-xxxx",  # Set the api_key
    base_url="https://api.deepseek.com")  # Set the request address


# Send a chat request
response = client.chat.completions.create(
    model="deepseek-v4-flash",  # Select the model
    messages=[
        {
            "role": "user",  # Role: user represents the user
            "content": "Please introduce EXAMPLE's Rookie Tutorial in one sentence."  # User input content
        }
    ]
)

# Print the full response (see what was returned)
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")

# Extract the AI's reply
ai_reply = response.choices[0].message.content
print("AI reply:")
print(ai_reply)

Run this program:

python first_api_call.py

If everything works normally, you will see output similar to this:

完整响应:
ChatCompletion(
  id='chatcmpl-xxx',
  choices=[Choice(finish_reason='stop', index=0, message=ChatMessage(content='Example是一个专注于提供编程技术教程的学习平台...', role='assistant'))],
  model='gpt-3.5-turbo',
  usage=CompletionUsage(completion_tokens=30, prompt_tokens=18, total_tokens=48)
)

==================================================

AI 回复:
Example是一个专注于提供编程技术教程的学习平台,涵盖多种编程语言和技术领域。

Congratulations! You have successfully called an AI API.

Request Parameters Explained

The OpenAI API has many parameters that can be adjusted. Here are the most commonly used ones:

Parameter NameTypeRequiredDescriptionDefault Value
modelstringYesThe model ID to use, such as gpt-3.5-turbo, gpt-4None
messagesarrayYesList of conversation history messagesNone
temperaturenumbernoSampling temperature, between 0-2. The higher it is, the more random; the lower it is, the more deterministic.1.0
max_tokensintegernoMaximum number of tokens to generateUnlimited
top_pnumbernoNucleus sampling parameter, between 0-11.0
stopstring/arraynoStop sequence; stop generation when these characters are encounterednull

Let's focus on temperature, which is the most commonly used adjustment parameter:

Examples

# File path: temperature_demo.py
# Demonstrate the effect of the temperature parameter on output

from openai import OpenAI

client = OpenAI(api_key="your-api-key-here")

def ask_ai(temperature_value: float, question: str) -> str:
    """Specify the temperature for the question"""
    response = client.chat.completions.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": question}],
        temperature=temperature_value,
        max_tokens=200
    )
    return response.choices[0].message.content

# Ask the same question three times with different temperature
question = "Give EXAMPLE a slogan"

print("--- temperature=0 (most deterministic, results are similar each time) ---")
print(ask_ai(0.0, question))
print("\n--- temperature=0.7 (balanced, recommended for daily use) ---")
print(ask_ai(0.7, question))
print("\n--- temperature=1.8 (most random, results vary greatly each time) ---")
print(ask_ai(1.8, question))

A rule of thumb:

  • temperature=0: factual answers, code generation, scenarios requiring precise results

  • temperature=0.7: everyday conversation, general tasks

  • temperature>1.0: creative writing, brainstorming


Anthropic Claude API

Claude is Anthropic's large language model, known for its safety and long-text processing capabilities.

Similarities and Differences with OpenAI API

Both can do text generation, but there are some differences in design philosophy:

Comparison itemOpenAI APIClaude API
Message structuresystem, user, assistant alternatesystem is set separately, user and assistant alternate
Context length16k-128k variesClaude 3 series supports 200k
Parameter designtemperature, top_p, etc.temperature, top_p, etc., similar concepts

Claude Messages API Structure

First install the Claude SDK:

pip install anthropic

Basic calls to the Claude API:

Examples

# File path: claude_first_call.py
# Basic Claude API call example

import anthropic

# Initialize the client
client = anthropic.Anthropic(
    api_key="your-api-key-here"  # Replace with your Claude API Key
)

# Send message request
response = client.messages.create(
    model="claude-3-sonnet-20240229",  # Select the Claude model
    max_tokens=1024,  # Claude requires max_tokens to be set
    messages=[
        {
            "role": "user",
            "content": "Please introduce EXAMPLE Rookie Tutorial in three sentences"
        }
    ]
)

# Print response
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")

# Extract reply
print("Claude's reply:")
print(response.content[0].text)

Similarly, we can replace it with a domestically compatible large model, such as DeepSeek, specifyingapi_key 、 base_url、modelThese three parameters:

Examples

import anthropic

# Initialize client
client = anthropic.Anthropic(
    api_key="sk-xxx",  # Replace with your API Key
    base_url="https://api.deepseek.com/anthropic"  # Base URL for Claude API
)

# Send message request
response = client.messages.create(
    model="deepseek-v4-flash",  # Choose Claude model
    max_tokens=1024,  # Claude requires max_tokens to be set
    messages=[
        {
            "role": "user",
            "content": "Please introduce the EXAMPLE tutorial in three sentences."
        }
    ]
)

# Print response
print("Full response:")
print(response)
print("\n" + "="*50 + "\n")

# Extract reply
print("Claude reply:")
print(response.content[0].thinking)

System Prompt Settings

In the Claude API, System Prompt is an independent parameter, not placed in the messages array:

Examples

# File path: claude_system_prompt.py
# Claude System Prompt usage example

import anthropic

client = anthropic.Anthropic(api_key="your-api-key-here")

response = client.messages.create(
    model="claude-3-sonnet-20240229",
    max_tokens=1024,
    # System Prompt is set here
    system="You are a professional technical documentation editor. Keep answers concise and accurate, and use lists where appropriate.",
    messages=[
        {
            "role": "user",
            "content": "What core knowledge points should be mastered when learning Python?"
        }
    ]
)

print(response.content[0].text)

Multi-turn Conversation Management

Single-turn conversations are simple, but truly useful applications need to "remember context."

Ways to Maintain Conversation History

The approach is simple:Store all previous conversations and send them to the AI with each request。

Examples

# File path: multi_turn_chat.py
# Multi-turn conversation management example

from openai import OpenAI

client = OpenAI(api_key="your-api-key-here")

class ChatBot:
    """A simple multi-turn conversation bot"""

    def __init__(self, system_prompt: str = "You are a helpful assistant"):
        self.messages = []  # Use a list to store conversation history
        # First add the system prompt
        self.messages.append({
            "role": "system",
            "content": system_prompt
        })

    def chat(self, user_input: str) -> str:
        """Send a message and get a reply"""
        # Add user message to history
        self.messages.append({
            "role": "user",
            "content": user_input
        })

        # Call the API
        response = client.chat.completions.create(
            model="gpt-3.5-turbo",
            messages=self.messages
        )

        # Get AI reply
        ai_reply = response.choices[0].message.content

        # Also add the AI reply to history
        self.messages.append({
            "role": "assistant",
            "content": ai_reply
        })

        return ai_reply

    def get_history(self) -> list:
        """Get conversation history"""
        return self.messages


# Usage example
if __name__ == "__main__":
    bot = ChatBot(system_prompt="You are a programming assistant specializing in answering questions about EXAMPLE and Python")

    print("=== Multi-turn conversation demo ===")

    # First round
    r1 = bot.chat("What is Python?")
    print(User: What is Python?)
    print("AI:", r1)
    print()

    # Second round (AI can remember the previous topic)
    r2 = bot.chat(What are its features?)
    print(User: What are its features?)
    print("AI:", r2)
    print()

    # Third round
    r3 = bot.chat(Are there Python tutorials on EXAMPLE?)
    print(User: Are there Python tutorials on EXAMPLE?)
    print("AI:", r3)

Key point: The conversation history is an ordered list, strictly alternating between user-assistant-user-assistant. Each request must send the full history so the AI can understand the context. History consumes tokens, so it needs to be cleaned up regularly.

Handling Context Window Overflow

Each model has a context window limit, e.g., gpt-3.5-turbo is 16k, gpt-4 is 8k or 32k.

When the conversation history is too long, an error is triggered.

Several common handling strategies:

Examples

# File path: context_management.py
# Example of context window management strategies

from openai import OpenAI

client = OpenAI(api_key="your-api-key-here")


class ChatBotWithContextLimit:
    """Chatbot with context limit"""

    def __init__(self, system_prompt: str, max_messages: int = 10):
        self.system_prompt = system_prompt
        self.max_messages = max_messages  # How many messages to keep at most
        self.messages = []

    def _trim_messages(self):
        """Trim message list, keep the most recent messages"""
        # Always keep the system prompt, then only keep the most recent max_messages
        if len(self.messages) > self.max_messages:
            # Keep the most recent N messages
            self.messages = self.messages[-self.max_messages:]

    def chat(self, user_input: str) -> str:
        # Add user message
        self.messages.append({
            "role": "user",
            "content": user_input
        })

        # Trim history
        self._trim_messages()

        # Construct full request (system prompt + history)
        full_messages = [
            {"role": "system", "content": self.system_prompt}
        ] + self.messages

        try:
            response = client.chat.completions.create(
                model="gpt-3.5-turbo",
                messages=full_messages
            )
            ai_reply = response.choices[0].message.content

            self.messages.append({
                "role": "assistant",
                "content": ai_reply
            })

            return ai_reply

        except Exception as e:
            # If it still errors, more aggressive trimming may be needed
            return f"Error: {str(e)}"


# Strategy 2: Use AI to summarize conversation history
def summarize_conversation(messages: list) -> str:
    """Use AI to summarize the conversation history, replacing the original history"""
    history_text = "\n".join([
        f"{m['role']}: {m['content']}"
        for m in messages
    ])

    prompt = f"""Please summarize the following conversation into a concise summary, preserving key information:

{history_text}

Summary: """


    response = client.chat.completions.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": prompt}],
        temperature=0
    )

    return response.choices[0].message.content


# Use summarization strategy
print("--- Conversation history summary example ---")
sample_history = [
    {"role": "user", "content": "What's my name?"},
    {"role": "assistant", "content": "Sorry, I don't know your name. You can tell me."},
    {"role": "user", "content": "My name is Xiao Ming, I'm learning Python at EXAMPLE."},
    {"role": "assistant", "content": "Hello Xiao Ming! Nice to meet you. EXAMPLE is a great learning platform."},
    {"role": "user", "content": "I want to learn about lists, can you teach me?"},
    {"role": "assistant", "content": "Of course! Python lists are..."},
]

summary = summarize_conversation(sample_history)
print("Conversation summary:", summary)

Summarize the methods of context management:

StrategyAdvantagesDisadvantagesApplicable scenarios
Keep only the most recent N entriesSimple, fastMay lose important informationShort conversations, casual chat scenarios
Use AI to summarize historyPreserves semantic informationRequires additional API callsLong conversations, many important pieces of information
Sliding windowBalances simplicity and completenessSlightly more complex to implementGeneral purpose for common scenarios

Streaming Output

When chatting with AI on a web page, you see text appear word by word — that is streaming output.

What is Streaming Output

  • Normal mode: AI returns the complete reply at once after generating it; you may have to wait several seconds or even tens of seconds.

  • Streaming mode: AI sends each token as soon as it is generated, providing a better user experience.

Technically, streaming output uses the SSE (Server-Sent Events) protocol, where the server continuously pushes data.

Implementing the Typewriter Effect

Examples

# File path: streaming_demo.py
# Streaming output demo (typewriter effect)

from openai import OpenAI

client = OpenAI(api_key="your-api-key-here")

print("=== Streaming Output Demo ===")
print("Q: Please write an introduction to EXAMPLE\n")
print("A:", end="", flush=True)

# Key parameter: stream=True
response = client.chat.completions.create(
    model="gpt-3.5-turbo",
    messages=[
        {"role": "user", "content": "Please write an introduction to the EXAMPLE tutorial"}
    ],
    stream=True  # Enable streaming output
)

collected_reply = []

for chunk in response:
    # Extract the content of the current chunk
    if chunk.choices[0].delta.content:
        content = chunk.choices[0].delta.content
        print(content, end="", flush=True)  # Print in real time without newlines
        collected_reply.append(content)

print()  # Output a newline at the end
print("\n" + "="*50)
print("Complete reply:", "".join(collected_reply))

Run this program, and you will see text appear word by word.

Claude's streaming output usage is similar:

Examples

# File path: claude_streaming.py
# Claude streaming output example

import anthropic

client = anthropic.Anthropic(api_key="your-api-key-here")

print("=== Claude Streaming Output ===")
print("Q: How to learn programming?\n")
print("A:", end="", flush=True)

with client.messages.stream(
    model="claude-3-sonnet-20240229",
    max_tokens=1024,
    messages=[{"role": "user", "content": "How to learn programming systematically?"}]
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

print()

User experience advice: As long as the scenario does not require "must wait for the complete result", prioritize streaming output. Letting users see progress can significantly improve the experience.


API Cost Calculation and Optimization

AI APIs are billed based on usage; the more you use, the more you spend.

Token Billing Principles

AI API is not billed by word count or number of conversations, but byToken,Tokenis the basic unit of AI text processing.

One token is approximately equal to:

  • English: 0.75 words (or 4 letters)

  • Chinese: 1-2 Chinese characters

For example:

  • "Hello, World" — approximately 3-4 tokens

  • "EXAMPLE Rookie Tutorial is great" — approximately 6-8 tokens

Billing method:

Prompt tokens(你发给 AI 的) + Completion tokens(AI 返回给你的)= 总 tokens

Different models have different prices, for example gpt-3.5-turbo is $0.0015/1k prompt tokens, $0.002/1k completion tokens.

How to Control API Costs

Examples

# File path: cost_tracking.py
# API cost tracking and control example

from openai import OpenAI

client = OpenAI(api_key="your-api-key-here")

# Model price list (example, actual prices please check official documentation)
# Unit: USD / 1k tokens
MODEL_PRICES = {
    "gpt-3.5-turbo": {
        "prompt": 0.0015,
        "completion": 0.002
    },
    "gpt-4": {
        "prompt": 0.03,
        "completion": 0.06
    }
}


def chat_with_cost_tracking(model: str, messages: list):
    """Chat with cost statistics"""
    response = client.chat.completions.create(
        model=model,
        messages=messages
    )

    # Get token usage from the response
    usage = response.usage
    prompt_tokens = usage.prompt_tokens
    completion_tokens = usage.completion_tokens
    total_tokens = usage.total_tokens

    # Calculate the cost
    price = MODEL_PRICES.get(model, {"prompt": 0, "completion": 0})
    prompt_cost = (prompt_tokens / 1000) * price["prompt"]
    completion_cost = (completion_tokens / 1000) * price["completion"]
    total_cost = prompt_cost + completion_cost

    return {
        "reply": response.choices[0].message.content,
        "usage": {
            "prompt_tokens": prompt_tokens,
            "completion_tokens": completion_tokens,
            "total_tokens": total_tokens
        },
        "cost": {
            "prompt_cost": prompt_cost,
            "completion_cost": completion_cost,
            "total_cost": total_cost
        }
    }


# Usage example
if __name__ == "__main__":
    result = chat_with_cost_tracking(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": "Introduce Python"}]
    )

    print("AI reply:", result["reply"])
    print()
    print("Usage statistics:")
    print(f"  Prompt tokens: {result['usage']['prompt_tokens']}")
    print(f"  Completion tokens: {result['usage']['completion_tokens']}")
    print(f"  Total tokens: {result['usage']['total_tokens']}")
    print()
    print("Cost statistics (USD):")
    print(f"  Prompt cost: ${result['cost']['prompt_cost']:.6f}")
    print(f"  Completion cost: ${result['cost']['completion_cost']:.6f}")
    print(f"  Total cost: ${result['cost']['total_cost']:.6f}")

Some practical tips for cost optimization:

TipHow to do itSavings ratio
Choose the right modelUse gpt-3.5-turbo for simple tasks, not gpt-490%+
Limit max_tokensSet a reasonable upper limit to avoid the AI saying too muchDepends on the situation
Trim contextKeep only necessary conversation history50%+
Keep prompts conciseDon't make the system prompt too longDepends on the situation
Result cachingReturn cached results directly for identical questionsDepends on the situation

Caching Strategies

For repeated questions, caching results can save a lot of money:

Examples

# File path: response_cache.py
# Simple response cache implementation

import hashlib
import json
from openai import OpenAI

client = OpenAI(api_key="your-api-key-here")


class SimpleCache:
    """Simple in-memory cache"""

    def __init__(self):
        self.cache = {}

    def _make_key(self, model: str, messages: list) -> str:
        """Generate cache key based on request content"""
        # Serialize key information and compute a hash
        key_data = {
            "model": model,
            "messages": messages
        }
        key_str = json.dumps(key_data, sort_keys=True)
        return hashlib.md5(key_str.encode()).hexdigest()

    def get(self, model: str, messages: list):
        """Get cache"""
        key = self._make_key(model, messages)
        return self.cache.get(key)

    def set(self, model: str, messages: list, value: str):
        """Set cache"""
        key = self._make_key(model, messages)
        self.cache[key] = value


# Chat function with cache
cache = SimpleCache()

def chat_with_cache(model: str, messages: list):
    # Check cache first
    cached = cache.get(model, messages)
    if cached:
        print("(Cache hit, return directly)")
        return cached

    # Cache miss, call API
    response = client.chat.completions.create(
        model=model,
        messages=messages
    )
    reply = response.choices[0].message.content

    # Store in cache
    cache.set(model, messages, reply)

    return reply


# Test
if __name__ == "__main__":
    messages = [{"role": "user", "content": "What kind of website is EXAMPLE?"}]

    print("First call...")
    r1 = chat_with_cache("gpt-3.5-turbo", messages)
    print(r1)
    print()

    print("Second call (same question)...")
    r2 = chat_with_cache("gpt-3.5-turbo", messages)
    print(r2)

Error Handling and Retry Mechanisms

Network requests can always fail, and APIs are no exception. Robust code must handle various exception scenarios.

Common Error Types

Error typeHTTP status codeCauseHandling method
Authentication failure401API Key incorrect or expiredCheck API Key
Insufficient quota402Account has no moneyRecharge or switch account
Limit exceeded429Request too fast or quota exhaustedWait a while and try again
Server error5xxAI server-side issueRetry

Exponential Backoff Retry

When encountering transient errors (such as 429, 5xx), the most common strategy is exponential backoff retry:

  • First failure → wait 1 second → retry

  • Second failure → wait 2 seconds → retry

  • Third failure → wait 4 seconds → retry

  • Fourth failure → wait 8 seconds → retry

Each wait time doubles, giving the service a chance to recover.

Examples

# File path: error_handling.py
# Error handling and exponential backoff retry

import time
from openai import OpenAI, APIError, RateLimitError, APIConnectionError

client = OpenAI(api_key="your-api-key-here")


def chat_with_retry(
    model: str,
    messages: list,
    max_retries: int = 3,
    initial_wait: float = 1.0
):
    """Chat function with retry mechanism"""

    retries = 0

    while retries <= max_retries:
        try:
            response = client.chat.completions.create(
                model=model,
                messages=messages
            )
            return {
                "success": True,
                "reply": response.choices[0].message.content
            }

        except RateLimitError as e:
            # Rate limit error: wait a while and try again
            retries += 1
            if retries > max_retries:
                return {
                    "success": False,
                    "error": f"Retry attempts exhausted: {str(e)}"
                }
            wait_time = initial_wait * (2 ** (retries - 1))
            print(f"Rate limit triggered, waiting {wait_time} seconds before retrying...")
            time.sleep(wait_time)

        except APIConnectionError as e:
            # Network connection error: retry
            retries += 1
            if retries > max_retries:
                return {
                    "success": False,
                    "error": f"Network connection failed: {str(e)}"
                }
            wait_time = initial_wait * (2 ** (retries - 1))
            print(f"Network error, waiting {wait_time} seconds before retrying...")
            time.sleep(wait_time)

        except APIError as e:
            # API error: retry depending on the situation
            if e.status_code and 500 <= e.status_code < 600:
                retries += 1
                if retries > max_retries:
                    return {
                        "success": False,
                        "error": f"Server error: {str(e)}"
                    }
                wait_time = initial_wait * (2 ** (retries - 1))
                print(f"Server error, waiting {wait_time} seconds before retrying...")
                time.sleep(wait_time)
            else:
                # Non-5xx errors, do not retry
                return {
                    "success": False,
                    "error": fAPI error: {str(e)}
                }

        except Exception as e:
            # Other errors, do not retry
            return {
                "success": False,
                "error": fUnknown error: {str(e)}
            }


# Usage example
if __name__ == "__main__":
    result = chat_with_retry(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": "Hello, EXAMPLE"}]
    )

    if result["success"]:
        print(Success:, result["reply"])
    else:
        print(Failure:, result["error"])

You can also use an existing library to simplify retry logic, such as tenacity, install it with the following command:

pip install tenacity

Examples

# File path: tenacity_demo.py
# Use tenacity library to simplify retry logic

from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
from openai import OpenAI, RateLimitError, APIError

client = OpenAI(api_key="your-api-key-here")


@retry(
    stop=stop_after_attempt(3),  # Retry up to 3 times
    wait=wait_exponential(multiplier=1, min=1, max=10),  # Exponential backoff: 1, 2, 4... up to 10 seconds
    retry=retry_if_exception_type((RateLimitError, APIError))  # Retry only on specific exceptions
)
def chat_with_tenacity(model: str, messages: list):
    """Use tenacity decorator to implement retry"""
    response = client.chat.completions.create(
        model=model,
        messages=messages
    )
    return response.choices[0].message.content


# Usage
try:
    reply = chat_with_tenacity(
        "gpt-3.5-turbo",
        [{"role": "user", "content": "Introduce Python"}]
    )
    print(reply)
except Exception as e:
    print(fFinal failure: {e})

Hands-on Project: Command-line AI Chat Tool

Now let's integrate what we've learned earlier and build a complete command line chat tool.

Examples

# File path: example_chatbot.py
# EXAMPLE AI Chatbot - Complete Command Line Tool

import sys
import json
import os
from datetime import datetime
from openai import OpenAI


class ExampleChatBot:
    """A fully functional command line chatbot"""

    def __init__(self, config_file: str = "config.json"):
        # Load configuration
        self.config = self._load_config(config_file)
        self.client = OpenAI(api_key=self.config["api_key"])

        # Conversation state
        self.messages = []
        self.total_tokens = 0
        self.total_cost = 0.0

        # Initialize system prompt
        self._init_system_prompt()

    def _load_config(self, config_file: str) -> dict:
        """Load configuration file"""
        default_config = {
            "api_key": "your-api-key-here",
            "model": "gpt-3.5-turbo",
            "system_prompt": "You are a helpful AI assistant powered by EXAMPLE.",
            "max_history": 20,
            "stream": True,
            "temperature": 0.7
        }

        if os.path.exists(config_file):
            with open(config_file, "r", encoding="utf-8") as f:
                user_config = json.load(f)
                default_config.update(user_config)

        return default_config

    def _save_config(self, config_file: str = "config.json"):
        """Save configuration file"""
        with open(config_file, "w", encoding="utf-8") as f:
            json.dump(self.config, f, indent=2, ensure_ascii=False)

    def _init_system_prompt(self):
        """Initialize system prompt"""
        self.messages = [
            {
                "role": "system",
                "content": self.config["system_prompt"]
            }
        ]

    def _trim_history(self):
        """Trim conversation history"""
        # Keep system prompt + the most recent N conversations
        max_messages = self.config["max_history"]
        if len(self.messages) > max_messages + 1:
            self.messages = [self.messages[0]] + self.messages[-max_messages:]

    def chat(self, user_input: str) -> str:
        """Send message and get reply"""
        # Add user message
        self.messages.append({
            "role": "user",
            "content": user_input
        })

        # Trim history
        self._trim_history()

        if self.config["stream"]:
            return self._chat_stream()
        else:
            return self._chat_normal()

    def _chat_normal(self) -> str:
        """Normal mode (non-streaming)"""
        response = self.client.chat.completions.create(
            model=self.config["model"],
            messages=self.messages,
            temperature=self.config["temperature"]
        )

        ai_reply = response.choices[0].message.content

        # Count usage
        if hasattr(response, "usage"):
            self.total_tokens += response.usage.total_tokens
            # Roughly estimate cost (gpt-3.5-turbo price)
            self.total_cost += (response.usage.total_tokens / 1000) * 0.002

        self.messages.append({
            "role": "assistant",
            "content": ai_reply
        })

        return ai_reply

    def _chat_stream(self) -> str:
        """Streaming mode"""
        response = self.client.chat.completions.create(
            model=self.config["model"],
            messages=self.messages,
            temperature=self.config["temperature"],
            stream=True
        )

        collected_chunks = []
        print("AI:", end="", flush=True)

        for chunk in response:
            if chunk.choices[0].delta.content:
                content = chunk.choices[0].delta.content
                print(content, end="", flush=True)
                collected_chunks.append(content)

        print()

        ai_reply = "".join(collected_chunks)

        self.messages.append({
            "role": "assistant",
            "content": ai_reply
        })

        # Token counting in streaming mode requires additional handling, simplified here
        # In real projects, you can call the usage API or estimate
        self.total_tokens += len(ai_reply) // 2  # Rough estimate

        return ai_reply

    def save_conversation(self, filename: str = None):
        """Save conversation history"""
        if not filename:
            timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
            filename = f"conversation_{timestamp}.json"

        data = {
            "model": self.config["model"],
            "total_tokens": self.total_tokens,
            "estimated_cost": self.total_cost,
            "messages": self.messages,
            "saved_at": datetime.now().isoformat()
        }

        with open(filename, "w", encoding="utf-8") as f:
            json.dump(data, f, indent=2, ensure_ascii=False)

        return filename

    def clear_history(self):
        """Clear conversation history"""
        self._init_system_prompt()
        self.total_tokens = 0
        self.total_cost = 0.0

    def show_stats(self):
        """Display statistics"""
        print("\n" + "="*40)
        print(Conversation Statistics)
        print("="*40)
        print(fModel: {self.config['model']})
        print(fMessage count: {len(self.messages) - 1})
        print(fTotal tokens: {self.total_tokens})
        print(fEstimated cost: ${self.total_cost:.4f})
        print("="*40 + "\n")


def print_help():
    """Print help information"""
    print("\n" + "="*40)
    print(EXAMPLE AI Chat Tool - Command List)
    print("="*40)
    print(/help - Show help)
    print(/clear - Clear conversation history)
    print(/stats - Show statistics)
    print(/save - Save conversation history)
    print(/model - Switch model)
    print(/temp - Adjust temperature)
    print(/system - Modify system prompt)
    print(/quit - Exit program)
    print("="*40 + "\n")


def main():
    print("="*40)
    print(EXAMPLE AI Chat Tool)
    print("="*40)
    print(Enter /help to view commands, or type text directly to start chatting\n")

    bot = ExampleChatBot()

    while True:
        try:
            user_input = input(You:).strip()

            if not user_input:
                continue

            # Handle commands
            if user_input.startswith("/"):
                cmd = user_input.lower()

                if cmd in ["/quit", "/exit", "/q"]:
                    print(Goodbye!)
                    break

                elif cmd == "/help":
                    print_help()

                elif cmd == "/clear":
                    bot.clear_history()
                    print(Conversation history cleared)

                elif cmd == "/stats":
                    bot.show_stats()

                elif cmd == "/save":
                    filename = bot.save_conversation()
                    print(fConversation saved to: {filename})

                elif cmd.startswith("/model "):
                    model = user_input[7:].strip()
                    bot.config["model"] = model
                    bot._save_config()
                    print(fModel switched to: {model})

                elif cmd.startswith("/temp "):
                    try:
                        temp = float(user_input[6:].strip())
                        if 0 <= temp <= 2:
                            bot.config["temperature"] = temp
                            bot._save_config()
                            print(fTemperature set to: {temp})
                        else:
                            print(Temperature must be between 0-2)
                    except ValueError:
                        print(Please enter a valid number)

                elif cmd.startswith("/system "):
                    system_prompt = user_input[8:].strip()
                    bot.config["system_prompt"] = system_prompt
                    bot.clear_history()
                    bot._save_config()
                    print(System prompt updated, conversation reset)

                else:
                    print(Unknown command, enter /help for help)

            else:
                # Regular chat
                if not bot.config["stream"]:
                    print("AI:", end="")

                bot.chat(user_input)
                print()

        except KeyboardInterrupt:
            print("\nEnter /quit to exit, or continue chatting)

        except Exception as e:
            print(fError: {e})


if __name__ == "__main__":
    main()

Create a configuration file:

{
    "api_key": "your-api-key-here",
    "model": "gpt-3.5-turbo",
    "system_prompt": "你是一个乐于助人的 AI 助手,由 EXAMPLE 提供技术支持。",
    "max_history": 20,
    "stream": true,
    "temperature": 0.7
}

Now you can run this chat tool:

python example_chatbot.py

You will see:

========================================
  EXAMPLE AI 聊天工具
========================================
输入 /help 查看命令,直接输入文字开始聊天

你:你好
AI:你好!很高兴见到你。我是由 EXAMPLE 提供技术支持的 AI 助手,有什么可以帮你的吗?

你:/stats

========================================
对话统计
========================================
模型:gpt-3.5-turbo
消息数:2
总 tokens:50
估算费用:$0.0001
========================================

你:/quit
再见!
Other extensions