Ollama Multi-Model Collaboration and Agent Workflow
A single model doesn't have to do everything: let small models handle routing, specialized models handle tasks, and connected models search for information. Use structured output as the scheduling protocol to combine fast and cost-effective intelligent applications.
This chapter implements a multi-model routing system and upgrades it into an Agent workflow with internet access capability.
Architecture Design: Small Model Routing + Expert Execution
The core idea is division of labor: determining "what the user wants to do" is simple and can be handled by the smallest model; the actual tasks are then passed to the appropriate specialized models.
This design has three advantages: small model routing takes only tens of milliseconds, so large models no longer handle low-value requests; the parameters of each expert model can be independently adjusted as needed; new task types only require adding a branch without affecting existing chains.
Implementation 1: Structured Output for Intent Routing
The key to routing is making classification results "machine-readable"; structured output serves as the scheduling protocol here.
Example
from ollama import chat
from pydantic import BaseModel, Field
# Structured definition of routing results
class Route(BaseModel):
category: str = Field(description=Question category: chat / coding / search)
reason: str = Field(description=One-sentence classification basis)
def classify(question: str) -> Route:
response = chat(
model='qwen3.5:4b', # A small model is sufficient for routing
messages=[{'role': 'user', 'content': question}],
format=Route.model_json_schema(),
options={'temperature': 0}, # Classification must be stable
)
return Route.model_validate_json(response.message.content)
# Quick self-test
for q in ['How to deduplicate a Python list?', 'What AI news is there today?', 'Tell a joke']:
r = classify(q)
print(r.category, '|', r.reason)
$ python router.py coding | 询问 Python 列表去重的编程方法 search | 询问当天的最新新闻,需要联网 chat | 闲聊类请求,直接回答即可
Implementation 2: Three Expert Processors
One processor function per path, with a unified signature for easy dispatch.
Example
from ollama import chat, web_search, web_fetch
def handle_chat(question):
"""Daily Q&A: answered directly by the small model"""
r = chat(model='qwen3.5:4b',
messages=[{'role': 'user', 'content': question}])
return r.message.content
def handle_coding(question):
"""Programming tasks: handed to a dedicated code model"""
r = chat(model='example-coder',
messages=[{'role': 'user', 'content': question}],
options={'temperature': 0.2, 'num_ctx': 65536})
return r.message.content
def handle_search(question):
"""Web-related tasks: Agent loop of small model + search tool"""
tools = {'web_search': web_search, 'web_fetch': web_fetch}
messages = [{'role': 'user', 'content': question}]
while True:
r = chat(model='qwen3.5:4b', messages=messages,
tools=[web_search, web_fetch])
messages.append(r.message)
if not r.message.tool_calls:
return r.message.content
for call in r.message.tool_calls:
func = tools.get(call.function.name)
# Truncate tool results to protect the context window
result = str(func(**call.function.arguments))[:2000]
messages.append({'role': 'tool',
'tool_name': call.function.name,
'content': result})
HANDLERS = {'chat': handle_chat,
'coding': handle_coding,
'search': handle_search}
def ask(question):
route = classify(question)
print(f'[Route -> {route.category}] {route.reason}')
return HANDLERS[route.category](question)
The router itself does not generate answers, so its hallucination risk is limited to "misclassification"; the cost of misclassification is just one extra correction and does not pollute the final answer.
Implementation 3: Local and Cloud Hybrid Orchestration
Each path can have a "cloud upgrade tier": for tasks that local models cannot handle, simply change the model name in the same code structure.
Example
def ask_with_fallback(question, force_cloud=False):
"""Scheduling with cloud fallback: local first, complex tasks upgraded to cloud"""
route = classify(question).category
if force_cloud or route == 'coding':
# Heavy tasks such as coding go directly to cloud flagship specs
model = 'qwen3.5:cloud'
else:
model = 'qwen3.5:4b'
print(f'[Route -> {route}] Using model {model}')
r = chat(model=model,
messages=[{'role': 'user', 'content': question}])
return r.message.content
Note a capability difference: Cloud models do not yet support structured output, so the router must be handled by a local model, which fits perfectly with the "small model for routing" design.
Comprehensive Evaluation of Cost and Latency
The cost-benefit of a multi-model strategy must be calculated across three dimensions simultaneously:
| Dimension | Local small model | Local large model | Cloud model |
|---|---|---|---|
| Time to first token | Lowest (millisecond-level loading) | Medium | Includes network round trip |
| Marginal cost | Approximately electricity cost | Electricity + VRAM usage | Billed by usage |
| Answer quality | Sufficient for simple tasks | Close to flagship | Flagship-level |
| Concurrency capability | Limited by the local machine | Limited by VRAM | Best elasticity |
This yields a practical baseline strategy: use local small models as the foundation for routing and lightweight tasks; assign quality-sensitive tasks to local large models; switch heavy-load or oversized tasks to the cloud; the model names of all three categories go through the configuration file and can be adjusted anytime based on retirement announcements or hardware changes.
Other Extensions