Ollama FAQ

This chapter condenses troubleshooting knowledge into a handy reference manual: start with log analysis, triage by symptom, and cover six major categories of problems - insufficient VRAM, download failures, port conflicts, GPU anomalies, overload, and upgrade compatibility.


The Universal First Step: Check the Logs First

Ollama's error messages are often very brief; the real root cause is almost always written in the logs.

EnvironmentView Command / Location
macOScat ~/.ollama/logs/server.log
Linux(systemd)journalctl -u ollama --no-pager --follow --pager-end
Windows%LOCALAPPDATA%\Ollama\server.log
Dockerdocker logs <container_name>
Manual serveLogs are printed directly in the current terminal

When the default logs lack detail, enable debug mode and restart the service to see key processes such as GPU discovery and VRAM estimation:

# 调试模式启动
OLLAMA_DEBUG=1 ollama serve

When encountering an anomaly, first triage using the navigation diagram below, then jump to the corresponding section:

排障导航:按症状分流四大类问题


Insufficient VRAM / Memory (OOM)

This is the most common runtime problem, typically manifesting as the process disappearing mid-generation, the system freezing, or memory allocation failures in the logs.

SymptomCauseCountermeasure
Process killed when loading a large modelInsufficient physical memorySwitch to a smaller tag; or use create to quantize and reduce the size
Crash during generationContext accumulation causes KV cache to expandLower num_ctx; enable Flash Attention and KV cache quantization
Collective failure under multiple requestsNUM_PARALLEL is too large, VRAM usage multipliesSet it back to 1, then gradually increase it while observing ollama ps
Model runs but is offloaded to CPUVRAM is slightly less than requiredClear background VRAM usage; drop down one parameter size level

The diagnostic command trio: ollama ps to check model occupancy and split ratio, nvidia-smi (or rocm-smi) to check real-time VRAM, and free -h to check system memory.


Model Download Failures and Resumable Interruptions

Download problems almost always lie in the network path; the good news is that Ollama supports layer-by-layer resumable downloads, making retries very cheap.

SymptomCauseCountermeasure
Downloads repeatedly interruptedUnstable network to the source serverRe-run the same command to resume; configure HTTPS_PROXY if necessary
Still fails after configuring proxyHTTP_PROXY was mistakenly setKeep only HTTPS_PROXY, delete HTTP_PROXY
Proxy certificate errorSelf-signed certificate is not trustedInstall the CA certificate as a system certificate
Disk space insufficient warningTarget partition capacity is not enoughClean up space, or use OLLAMA_MODELS to migrate the storage path

Port Conflicts and Connection Refused

When the client reports "connection refused", the essence is that no service is responding on port 11434.

SymptomCauseCountermeasure
Local curl connection refusedOllama service is not startedOpen the app on macOS / Windows; run sudo systemctl start ollama on Linux
Logs report port is already in use11434 is taken by another programStop the occupying process, or start on a different port: OLLAMA_HOST=127.0.0.1:11435
Remote machines cannot connectService is only listening on the loopback addressSet OLLAMA_HOST=0.0.0.0 on the server side and allow it through the firewall
Linux manual install exits immediately after startup/tmp is mounted with noexec, inference library cannot be loadedSet OLLAMA_TMPDIR to point to an executable directory

GPU Not Used / Discovery Failure

The symptom is ollama ps showing 100% CPU, or logs showing GPU discovery failure. NVIDIA and AMD have different troubleshooting paths.

NVIDIA Specifics

First check the driver and passthrough environment, then handle according to the log error codes:

Check ItemMethod
Driver versionnvidia-smi outputs normally; upgrade to the latest driver
Container passthroughIf docker run --gpus all ubuntu nvidia-smi fails, it's a Container Toolkit issue
Driver module anomalyReload with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm and retry
Fails after resume from suspendAfter Linux sleep/wake, reload the nvidia_uvm module
Docker loses GPU after running for a long timeAdd native.cgroupdriver=cgroupfs to /etc/docker/daemon.json and restart Docker

The CUDA error codes in the logs are worth memorizing: 3 means not initialized, 46 means device unavailable, 100 means no device, 999 means unknown error - most point to driver or passthrough configuration.

AMD Specifics

SymptomCauseCountermeasure
Completely undetected on LinuxUser is not in the video / render groupAdd the running user to the corresponding group and restart the service
Logs show discovery timeoutROCm driver is too old (below v7)Upgrade ROCm v7 with amdgpu-install and restart
Old cards are not supportedArchitecture is not in the official supported listTry HSA_OVERRIDE_GFX_VERSION to specify a similar architecture
Some old cards have no ROCm on WindowsDriver stack limitationUse the default Vulkan path; check GGML_VK_VISIBLE_DEVICES if abnormal

Systematic Troubleshooting for 503 Overload and Slow Response

Link the methods from previous chapters into a fixed procedure; within six steps you can locate the vast majority of performance problems:

Step one, use ollama ps to confirm whether PROCESSOR is 100% GPU; if splitting appears, go back to VRAM optimization first.

Step two, check the logs for records of queueing, unloading, or discovery failures.

Step three, receiving 503 means the queue is full; check the concurrent request volume and OLLAMA_MAX_QUEUE, and evaluate whether to scale up.

Step four, in multi-user shared scenarios, verify the multiplicative relationship between OLLAMA_NUM_PARALLEL and VRAM.

Step five, check the keep_alive strategy: frequent unloading and reloading will cause first-token latency to spike.

Step six, use the API's usage field to calculate token/s, and compare with the historical baseline to quantify the effect of each adjustment.


Post-Upgrade Compatibility Issues and Rollback

When behavior is abnormal after an upgrade, three actions resolve most cases:

# 1. Linux 手动升级前先清理旧库,避免新旧混用
sudo rm -rf /usr/lib/ollama

# 2. 回退到指定版本
curl -fsSL https://ollama.com/install.sh | OLLAMA_VERSION=0.5.7 sh

# 3. AMD 显卡同步升级 ROCm v7 驱动后重启

Garbled blocks appearing in the Windows terminal are an old terminal font issue, unrelated to the model; just switch to Windows Terminal.

Other Extensions