Ollama REST API Programming Access

After installation, Ollama provides a complete REST API on local port 11434, covering text generation, chat, vector embeddings, and model management, as well as two compatible protocols: OpenAI and Anthropic.

This article explains the core APIs one by one: request structure, streaming processing, usage statistics, and error handling, allowing any language to integrate with local models.


Two Base URLs and Authentication Methods

Ollama's API has two endpoints: local and cloud, with different authentication rules.

Ollama API 全景:调用方、服务端点、本地模型与 Cloud

EndpointBase URLAuthentication
Local Servicehttp://localhost:11434/apiNo authentication required
Local Service (Compatible Protocol)http://localhost:11434/v1No authentication required, api_key can be filled arbitrarily
Ollama Cloudhttps://ollama.com/apiAuthorization: Bearer + API Key

Local services require no authentication by default and work out of the box; when directly connecting to Ollama's cloud API at ollama.com, first create an API Key on the official website settings page and include it in the request header.


Text Generation: /api/generate

generate is the most basic single-turn generation API; just pass the model name and prompt:

Example

curl http://localhost:11434/api/generate -d '{
  "model": "qwen3.5",
"prompt": "Introduce Python in one sentence",
  "stream": false
}'

When stream is set to false, a single JSON object is returned:

{
  "model": "qwen3.5",
  "created_at": "2026-08-29T03:20:00.499127Z",
  "response": "Example是一个面向编程初学者的免费中文教程网站。",
  "done": true,
  "total_duration": 10706818083,
  "load_duration": 6338219291,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 130079000,
  "eval_count": 42,
  "eval_duration": 4232710000
}

Common parameters are as follows:

ParameterDescription
modelRequired, model name
promptPrompt; can be left empty to preload the model
streamDefault true returns streaming; when false, returns the complete result at once
systemTemporarily specify a system prompt, overriding the setting in Modelfile
optionsGeneration parameters object, e.g., temperature, num_ctx (same as Modelfile parameters)
format"json" or a full JSON Schema, forces structured output
keep_aliveDuration for the model to stay resident after the request, e.g., "10m", -1 for permanent residency, 0 for immediate unload
rawWhen true, skips template assembly and directly uses the full prompt you provide
suffixText to be appended after the model output, for text completion scenarios
imagesBase64 image array, used with multimodal models

Two practical tips: sending an empty prompt can preload the model into memory, eliminating the loading wait for the first request; setting keep_alive to 0 along with an empty prompt can immediately unload the model and release VRAM.

Example

# Preload model (no content generated)
curl http://localhost:11434/api/generate -d '{"model": "qwen3.5"}'

# Immediately unload model, release VRAM
curl http://localhost:11434/api/generate -d '{"model": "qwen3.5", "keep_alive": 0}'

Chat Generation: /api/chat

chat is a multi-turn conversation API; the message history is maintained by the caller, which is the essential difference from generate.

Example

curl http://localhost:11434/api/chat -d '{
  "model": "qwen3.5",
  "messages": [
{ "role": "user", "content": "What is Python?" },
{ "role": "assistant", "content": "A Chinese tutorial website for beginners." },
{ "role": "user", "content": "Is it free? Answer in one sentence" }
  ],
  "stream": false
}'

Fields supported by the message object:

FieldDescription
roleFour roles: system / user / assistant / tool
contentMessage content
imagesOptional, base64 image list (multimodal models)
thinkingOptional, the reasoning process of thinking models (used with the think parameter)
tool_callsOptional, the list of tools that the model requests to call (see the tool calling section for hands-on practice)

The correct way for multi-turn conversation: append the assistant message returned by the model in each round back to the messages array, then send it with the new question; the model can then "remember" the full context.

Since messages are fully managed by you, the history can be persisted to a database, compressed via summarization, or restored across sessions — the problem mentioned in Part 5 of "command line exits lose memory" now has an engineering solution here.


Vector Embeddings: /api/embed

The embed API converts text into vectors, serving as the raw material workshop for RAG and semantic search.

Example

# Single text
curl http://localhost:11434/api/embed -d '{
  "model": "embeddinggemma",
"input": "EXAMPLE is a programming tutorial website"
}'


# Batch: pass an array to input, returns multiple vectors at once
curl http://localhost:11434/api/embed -d '{
  "model": "embeddinggemma",
"input": ["first text", "second text", "third text"]
}'

The embeddings in the response is a 2D array, with each vector already L2-normalized (unit length), so cosine similarity can be directly used for comparison.

Two noteworthy parameters: dimensions can specify the output vector dimension (when supported by the model, reducing dimensions saves storage); truncate defaults to true and automatically truncates overly long text, and setting it to false will raise an error when the text is too long.


Streaming vs Non-streaming: How to Handle Responses

Generation APIs return streaming by default, in NDJSON format: each line is an independent JSON object.

{"model":"qwen3.5","created_at":"...","response":"EXAMPLE","done":false}
{"model":"qwen3.5","created_at":"...","response":"(Example)","done":false}
{"model":"qwen3.5","created_at":"...","response":"是编程初学者的入门网站。","done":false}
{"model":"qwen3.5","created_at":"...","response":"","done":true,"done_reason":"stop"}

The client reads line by line and concatenates the response fields to form the complete answer; the concatenation process can be rendered in real time on the interface.

Trade-offs between the two modes:

ModeAdvantageUse case
Streaming (default)Low time-to-first-token, can display in real timeChat interfaces, long-text generation
Non-streaming stream:falseGet the complete result at once, simple to processBatch processing, structured output, script calls

In streaming mode, errors may occur midway through the output; at that point the HTTP status code can no longer be changed. The error will appear as a line {"error": "..."} at the end of the data stream, and the client should check this line when parsing.


Usage Statistics and Error Handling

Each generation response ends with performance statistics fields, providing first-hand data for evaluating inference overhead.

FieldMeaning
total_durationTotal request time (nanoseconds)
load_durationModel load time (noticeable on first request)
prompt_eval_countNumber of tokens consumed by input
prompt_eval_durationInput processing time
eval_countNumber of tokens generated by output
eval_durationOutput generation time

The formula for generation speed (token/s): eval_count / eval_duration * 10^9; all time units are nanoseconds.

In terms of error handling, the API uses standard HTTP status codes to express results:

Status codeMeaning
200Success
400Request error (missing parameters, invalid JSON, etc.)
404Model does not exist (run ollama pull first or check the name)
429Rate limited due to too frequent requests
500Internal server error
502Gateway error (e.g., cloud model unreachable)

The error response body is fixed to the {"error": "error description"} structure; clients can simply extract the error field uniformly.


Model Management API Overview

The API provides a counterpart for every model management operation available in the command line, suitable for building admin dashboards or automation scripts.

EndpointMethodPurpose
/api/tagsGETList local models (including parameter count, quantization level)
/api/showPOSTView model details, capabilities, template
/api/pullPOSTPull a model, returns download progress streaming
/api/pushPOSTPush model to model library
/api/copyPOSTCopy model (source / destination)
/api/deleteDELETEDelete a model
/api/createPOSTCreate a model (including quantization, corresponding to Modelfile capabilities)
/api/psGETView currently loaded models and VRAM usage
/api/versionGETQuery Ollama version

Taking /api/tags as an example, the response contains specification details for each model:

{
  "models": [
    {
      "name": "qwen3.5:latest",
      "size": 6600000000,
      "details": {
        "family": "qwen3.5",
        "parameter_size": "9B",
        "quantization_level": "Q4_K_M"
      }
    }
  ]
}

Fields such as quantization_level and parameter_size can be directly used to build a model selector interface.


OpenAI Compatible Interface

For existing OpenAI applications to switch to local models, usually you only need to change one base_url.

Example

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5",
"messages": [{ "role": "user", "content": "Introduce Python in one sentence" }]
  }'

Compatibility covers five endpoints:

EndpointDescription
/v1/chat/completionsChat completions, supporting streaming, vision, tools, and JSON mode
/v1/completionsText completions
/v1/responsesResponses API (non-stateful mode)
/v1/embeddingsVector embeddings
/v1/modelsModel list

Pay attention to two common differences:

First, the OpenAI protocol has no num_ctx concept. When you need to change the context, first create a model with PARAMETER num_ctx and then call it under a new name (you can use the Modelfile from Part 7 or borrow a name with cp).

Second, some fields are not yet supported: tool_choice, logit_bias, n, logprobs, etc. Before migrating, it is recommended to check the manual's support list.


Anthropic Compatible Interface

Ollama also provides an Anthropic Messages API compatibility layer, allowing tools such as Claude Code to directly use local models.

Example

curl http://localhost:11434/v1/messages \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "qwen3.5",
    "max_tokens": 1024,
"messages": [{ "role": "user", "content": "Introduce Python in one sentence" }]
  }'

When integrating tools like Claude Code, you only need to set two environment variables:

Example

export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_BASE_URL=http://localhost:11434
claude --model qwen3.5

Note the capability boundaries: Anthropic features such as forced tool_choice, prompt caching, batch interface, and PDF input are not yet supported; token count is approximate.

Other extensions