Ollama Python Usage

Ollama provides a Python SDK that allows us to interact with locally running models in a Python environment.

With Ollama's Python SDK, you can easily integrate natural language processing tasks into Python projects, perform various operations such as text generation, conversation generation, and model management, without manually invoking the command line.

Install Python SDK

First, we need to install Ollama's Python SDK.

You can install it using pip:

pip install ollama

Make sure your environment has Python 3.x installed, and that your network environment can access the Ollama local service.

Start Local Service

Before using the Python SDK, make sure the Ollama local service has been started.

You can start it using the command line tool:

ollama serve

After starting the local service, the Python SDK will communicate with the local service to perform tasks such as model inference.

Using Ollama's Python SDK for Inference

After installing the SDK and starting the local service, we can interact with Ollama through Python code.

First, import chat and ChatResponse from the ollama library:

from ollama import chat
from ollama import ChatResponse

Through the Python SDK, you can send requests to a specified model to generate text or conversations:

Example

from ollama import chat
from ollama import ChatResponse

response: ChatResponse = chat(model='qwen3.5', messages=[
  {
    'role': 'user',
    'content': 'Who are you?',
  },
])
# Print response content
print(response['message']['content'])

# Or directly access the fields of the response object
#print(response.message.content)

Executing the above code, the output is:

你好!我是 Qwen3.5,是通义千问系列中最新推出的大语言模型。我具备强大的语言理解与生成能力,能协助你完成多种任务!

The Ollama SDK also supports streaming responses. When sending a request, we can setstream=Trueto enable response streaming.

Example

from ollama import chat

stream = chat(
    model='qwen3.5',
    messages=[{'role': 'user', 'content': 'Who are you?'}],
    stream=True,
)

# Print response content chunk by chunk
for chunk in stream:
    print(chunk['message']['content'], end='', flush=True)

Custom Client

You can also create a custom client to further control request configuration, such as setting custom headers or specifying the URL of the local service.

Create a Custom Client

Through Client, you can customize request settings (such as request headers, URL, etc.) and send requests.

Example

from ollama import Client

client = Client(
    host='http://localhost:11434',
    headers={'x-some-header': 'some-value'}
)

response = client.chat(model='qwen3.5', messages=[
    {
        'role': 'user',
        'content': 'Who are you?',
    },
])
print(response['message']['content'])

Asynchronous Client

If you want to execute requests asynchronously, you can use the AsyncClient class, which is suitable for scenarios requiring concurrency.

Example

import asyncio
from ollama import AsyncClient

async def chat():
    message = {'role': 'user', 'content': 'Who are you?'}
    response = await AsyncClient().chat(model='qwen3.5', messages=[message])
    print(response['message']['content'])

asyncio.run(chat())
different

The async client supports the same functionality as traditional synchronous requests; the only difference is that requests are executed asynchronously, which can improve performance, especially in high-concurrency scenarios.

Asynchronous Streaming Response

If you need to process streaming responses asynchronously, you can setstream=Trueit to an asynchronous generator to implement this.

Example

import asyncio
from ollama import AsyncClient

async def chat():
    message = {'role': 'user', 'content': 'Who are you?'}
    async for part in await AsyncClient().chat(model='qwen3.5', messages=[message], stream=True):
        print(part['message']['content'], end='', flush=True)

asyncio.run(chat())

Here, the response is returned asynchronously in parts, and each part can be processed immediately.


Common API Methods

The Ollama Python SDK provides some common API methods for operating and managing models.

1. chat method

Performs conversation generation with the model, sending user messages and getting model responses:

ollama.chat(model='llama3.2', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}])

2. generate method

Used for text generation tasks. Similar to the chat method, but it only requires a prompt parameter:

ollama.generate(model='llama3.2', prompt='Why is the sky blue?')

3. list method

Lists all available models:

ollama.list()

4. show method

Shows detailed information of the specified model:

ollama.show('llama3.2')

5. create method

Creates a new model from an existing model:

ollama.create(model='example', from_='llama3.2', system="You are Mario from Super Mario Bros.")

6. copy method

Copies a model to another location:

ollama.copy('llama3.2', 'user/llama3.2')

7. delete method

Deletes the specified model:

ollama.delete('llama3.2')

8. pull method

Pulls a model from a remote repository:

ollama.pull('llama3.2')

9. push method

Pushes a local model to a remote repository:

ollama.push('user/llama3.2')

10. embed method

Generates text embeddings:

ollama.embed(model='llama3.2', input='The sky is blue because of rayleigh scattering')

11. ps method

Views the list of running models:

ollama.ps()

Error Handling

The Ollama SDK throws errors when a request fails or when problems occur during response streaming.

We can use try-except statements to catch these errors and handle them as needed.

Example

model = 'does-not-yet-exist'

try:
    response = ollama.chat(model)
except ollama.ResponseError as e:
    print('Error:', e.error)
    if e.status_code == 404:
        ollama.pull(model)

In the above example, if the model does-not-yet-exist does not exist, a ResponseError is thrown. After catching it, you can choose to pull the model or perform other processing.

Other Extensions