Ollama Model Management

In this chapter, we will cover Ollama's model management commands: pulling, viewing, copying, deleting, and how models reside in memory.


Pull Model: ollama pull

The pull command only downloads the model without starting a conversation, making it suitable for preparing models in advance.

Download the default tag (latest):

ollama pull qwen3.5

Download a specific specification tag:

ollama pull qwen3.5:4b

After execution, you will see the standard download progress output:

$ ollama pull qwen3.5:4b
pulling manifest
pulling 87497c5f0e8f... 100% |████████████████████| 3.4GB
verifying sha256 digest
writing manifest
success

If a download is interrupted, you don't need to start over. Ollama verifies and stores data by layer; re-running the same command will automatically resume the unfinished layers.

It should also be noted that ollama run has built-in pull logic: the first time you run a model, downloading and loading are completed automatically, so in daily use you often don't need to pull separately.

Different specification tags correspond to different official pre-built model files; you cannot specify additional quantization precision when pulling. If you want to compress a large model into a smaller size yourself, you need to use the create command for quantization, which is covered in the Modelfile customization chapter.


View Local Models: list and show

list lists all local models and their disk usage; it is the most commonly used inventory command.

List local models:

ollama list

For example:

$ ollama list
NAME               ID            SIZE      MODIFIED
qwen3.5:0.8b       9e3f4b3a1c2d  1.0GB     5 minutes ago
qwen3.5:4b         b1c7d9e2f3a4  3.4GB     2 days ago
qwen3.5:latest     a8b2c9d3e4f5  6.6GB     2 hours ago

The four columns are the model name, version fingerprint, disk usage, and last modified time. When disk space is running low, checking the SIZE column is the most direct approach.

To look deeper into a model's details, use the show command:

# 查看模型详情:能力、参数、系统提示词、许可证等
ollama show qwen3.5

show can answer several typical questions: what capabilities the model supports (chat, vision, tool calling), what the default generation parameters are, what the system prompt says, and how large the context window is.

It also has an advanced use: exporting the model's complete configuration as a Modelfile:

Example

# Export the model recipe, which you can modify to rebuild your own model
ollama show --modelfile qwen3.5

This "view the recipe - modify - rebuild" workflow will be fully covered in the Modelfile customization chapter.


View Running Models: ollama ps

list looks at disk, while ps looks at memory: what has been loaded, how much it occupies, and when it will be unloaded.

Example

# List models currently loaded in memory/VRAM
ollama ps
$ ollama ps
NAME             ID            SIZE      PROCESSOR    CONTEXT    UNTIL
qwen3.5:9b       a8b2c9d3e4f5  9.6GB     100% GPU     131072     4 minutes from now

The meaning of each column is as follows:

ColumnMeaning
NAME / IDModel name and version fingerprint, same as in list
SIZETotal memory currently occupied by the model (including RAM and VRAM)
PROCESSORLoading location: 100% GPU means fully on the graphics card; 100% CPU means fully in RAM; a ratio such as 48%/52% CPU/GPU indicates insufficient VRAM and the model has been split
CONTEXTThe currently allocated context window size
UNTILTime remaining before automatic unload (keep-alive countdown)

The PROCESSOR column is the first place to check when troubleshooting "why is the model slow": as long as it's not 100% GPU, it means the model isn't fully loaded into VRAM. Specific performance optimization methods are discussed in detail in the GPU and Performance chapter.


Copy and Rename: ollama cp

cp gives an existing model a new name, commonly used for renaming or keeping a snapshot.

Example

# Copy a newly named model
ollama cp qwen3.5:4b example-test

# Use it like a normal model
ollama run example-test

It has a very practical advanced trick: "borrowing a name" for a model.

Some tools hardcode model names (e.g., gpt-3.5-turbo). By using cp to copy a local model to that name, you can let such tools directly use the local model:

Example

# Copy the local model to the name expected by the tool
ollama cp qwen3.5 gpt-3.5-turbo

cp completes almost instantly because it only creates a new model name entry pointing to the original model files; it doesn't create a second multi-GB copy of data on disk.


Delete Models and Disk Cleanup: ollama rm

rm deletes local models and is the antidote to the "only ever grows, never shrinks" problem in the model library.

Example

# Delete the specified model
ollama rm qwen3.5:0.8b
$ ollama rm qwen3.5:0.8b
deleted 'qwen3.5:0.8b'

The recommended cleanup workflow is: first use ollama list to find models that are large and no longer used, then rm them one by one.

If multiple tags share the same underlying files, deleting one of the tags won't affect the normal use of the other tags.

Make sure to also handle models whose names were borrowed with cp: after deleting the original model, the name-borrowed copy still exists and continues to occupy space. Don't let them become invisible disk black holes.


Stopping Models and the keep-alive Residency Mechanism

Let's first look at a model lifecycle diagram, and it will be clear at a glance which step each management command acts on.

Ollama 模型生命周期:下载、加载、驻留与卸载

The key mechanism is keep-alive: after a model conversation ends, the model isn't unloaded immediately, but by default stays resident in memory or VRAM for 5 minutes.

If another request comes within those 5 minutes, the response no longer waits for model loading; if there are no new requests after that time, the model is automatically unloaded and memory is released.

To make the model release resources immediately, use the stop command to unload it manually:

Example

# Immediately unload the specified model
ollama stop qwen3.5:4b

There are three ways to adjust the residency duration:

MethodHow to do itUse case
Change global defaultSet the environment variable OLLAMA_KEEP_ALIVE, e.g., 30m, 2hWhen the server wants a unified residency policy
Specify per requestPass the keep_alive parameter when calling the API; for example, -1 means it stays resident and is never unloadedWhen the application layer treats different models differently
Unload immediatelyollama stop <model name>Temporarily free VRAM for other tasks
Other Extensions