Ollama Model Library and Model Selection

This chapter will horizontally review how to read the Ollama model library, the positioning of mainstream model families, and the trade-offs of quantization precision, ultimately giving you a practical method for selecting models by task.

If the model is chosen correctly, half of the performance problems no longer exist.


How to Read the Model Library: Capability Tag System

Ollama model library (ollama.com/search) uses a tagging system to label each model's capabilities. The first step in selecting a model is to filter by tags.

TagMeaningTypical Use
visionSupports image inputImage Q&A, screenshot understanding, document recognition
toolsSupports tool callingAgent, function calling, automation
thinkingSupports thinking modeMathematics, logic, complex reasoning
embeddingText to vectorRAG, semantic search, knowledge base
cloudRuns in the cloud, no local compute requiredCall large models when local configuration is insufficient

Taking qwen3.5 used in this tutorial as an example, it carries the vision, tools, and thinking tags simultaneously in the model library, making it an "all-rounder". This is also why the entire book chooses it as the test model.

Each model page also lists all tags, download size, context window, input types, and update time. Spending one minute reading them before downloading can save a lot of trial-and-error time.


Overview of Mainstream Model Families

Ollama is not a model, but a platform for running various open-source models. Below are several of the most common families currently:

FamilyProducerPositioning & FeaturesSuitable Scenarios
Qwen (Tongyi Qianwen)AlibabaComplete range of sizes, strong Chinese capability, full multimodal and tool calling supportChinese tasks, general development, examples in this tutorial
DeepSeekDeepSeekThe reasoning series (R1) excels in chain-of-thoughtMathematics, logic, tasks requiring deep reasoning
GemmaGoogleExcellent performance in lightweight sizes, multimodalLow-spec devices, lightweight tasks
LlamaMetaBroadest ecosystem, abundant community resourcesGeneral tasks, English scenarios
GPT-OSSOpenAIOpenAI's open-source series, supports thinking levelsGeneral tasks, Agent experiments
MistralMistral AIRepresentative of the European camp, efficient small modelsLightweight deployment, multilingual

Model iteration speed is measured in weeks; specific available sizes and capabilities are subject to the real-time pages of the model library. The selection mindset outlives specific conclusions: first determine the task and capability tags, then look at the match between size and hardware.


Quantization Precision: Trade-off Between Size and Quality

Quantization is a technique that compresses model weights from high-precision decimals into low-bit-width storage. It is the key to "running large models on small GPUs".

For different quantized versions of the same model, the relationship between size and quality is roughly as follows:

PrecisionRelative SizeQuality PerformanceApplicable Scenarios
fp16 / fp32Baseline (largest)LosslessOriginal weights before training and quantization
q8_0About 1/2Almost losslessPreferred when VRAM is abundant
q4_K_MAbout 1/4Slight decrease, best cost-performanceChoice for the vast majority of local deployments
q4_K_SSlightly smaller than q4_K_MA bit lower than q4_K_MExtreme VRAM optimization

Models in the model library are already quantized by default (commonly q4_K_M), ready to use after download, no need to worry about it yourself.

To confirm the actual precision of a local model, you can query the quantization level using the show series commands or the API; if you want to compress an fp16 model into a small size yourself, just use the quantize parameter of the create command. The next Modelfile chapter has a complete demonstration.


Selecting Models by Task

Turn "model selection" into a classification problem: first determine the task, then compare with recommendations.

按任务选择模型的决策树

TaskRecommended ModelHardware ReferenceRemarks
Everyday conversation, writingqwen3.5:4b ~ 9bRuns with 16GB RAMBest cost-performance choice, full capability tags
Code generationqwen3-coder24GB or more VRAM recommendedRequires 64K context when used with coding tools
Mathematics, logical reasoningdeepseek-r1 / qwen3.5 thinking modeRegular configuration according to sizeMore stable but slower after enabling think
Image viewing, screenshot understandingqwen3.5 (native vision)Same as conversation sizePass image path directly on command line
RAG, semantic searchembeddinggemma / qwen3-embeddingSmall size, low requirementsEmbedding models cannot be used for chat

Embedding models and generative models are two different things: the former only converts text into vectors for retrieval and similarity computation, and does not output text answers. When building a knowledge base, you need both; don't mix them up.


Matching Size with Hardware

The elimination method is more useful than memorization: first use a rough rule to determine whether it can run, then fine-tune.

Rough rule: the model download size is approximately equal to the minimum RAM/VRAM requirement. For actual operation, reserve an additional 1-2GB overhead; the larger the context window, the more you need to add.

Your available RAM / VRAMMaximum size that can runRecommendation
8GBqwen3.5:0.8b ~ 2bExperience-oriented, accept the capability ceiling
16GBqwen3.5:4b ~ 9bSweet spot for daily use
8GB VRAM (dedicated GPU)qwen3.5:9bKeep 100% GPU for speed
24GB VRAMqwen3.5:27b / qwen3-coderAdvanced development and coding tasks
48GB VRAM or above35b+ and multi-GPU combinationsDefault context also increases automatically

Once it is running, useollama psVerification: If the PROCESSOR column shows 100% GPU, it means everything is in VRAM; if a CPU/GPU mixed ratio appears, it means VRAM is insufficient and the model is split, causing a noticeable speed drop. In this case, you should switch to a smaller size rather than tough it out.


Cloud Models and Retirement Mechanism

When local compute is insufficient but you still want to use very large models, Ollama provides a third path: cloud models with the cloud tag.

Examples

# Login to Ollama account (required for cloud models)
ollama signin

# Run cloud models like local models, no need to download weights
ollama run qwen3.5:cloud

Cloud models are used in exactly the same way as local models: the same commands, the same API, the same tool integrations, except that inference happens on Ollama's servers.

Note that cloud models have a retirement mechanism: the official team periodically takes down old models and provides recommended replacements, notifying users via email and official website announcements. Automated workflows that depend on a cloud model should pay attention to such notices.

Local models are completely unaffected by the retirement mechanism. Weights downloaded to disk can always be used, which is also the feature that makes pure local solutions most reassuring for teams.

Other Extensions