Ollama Related Commands
Ollama provides a variety of command-line tools (CLI) for users to interact with locally running models.
Basic format:
ollama <command> [args]
We can useollama --helpto see what commands are available:
Large language model runner Usage: ollama [flags] ollama [command] Available Commands: serve Start ollama create Create a model from a Modelfile show Show information for a model run Run a model stop Stop a running model pull Pull a model from a registry push Push a model to a registry list List models ps List running models cp Copy a model rm Remove a model help Help about any command Flags: -h, --help help for ollama -v, --version Show version information
1. Usage
ollama [flags]: Run ollama with flags.
ollama [command]: Run a specific ollama command.
2. Available Commands
- serve: Start the ollama service.
- create: Create a model from a Modelfile.
- show: Display detailed information about a model.
- run: Run a model.
- stop: Stop a running model.
- pull: Pull a model from a model registry.
- push: Push a model to a model registry.
- list: List all models.
- ps: List all running models.
- cp: Copy a model.
- rm: Delete a model.
- help: Get help information about any command.
3. Flags
- -h, --help: Display help information for ollama.
- -v, --version: Display version information.
Complete example:
| Command | Description | Example |
|---|---|---|
ollama run |
Run a model. If it does not exist, it will be pulled automatically. | ollama run llama3 |
ollama pull |
Pull a model. Download the model from the registry without running it. | ollama pull mistral |
ollama list |
List models. Display all locally downloaded models. | ollama list |
ollama rm |
Delete a model. Remove the local model to free up space. | ollama rm llama3 |
ollama cp |
Copy a model. Copy an existing model to a new name (for testing). | ollama cp llama3 my-model |
ollama create |
Create a model. Create a custom model from a Modelfile (advanced). | ollama create my-bot -f ./Modelfile |
ollama show |
Show information. View the model's metadata, parameters, or Modelfile. | ollama show --modelfile llama3 |
ollama ps |
View processes. Display currently running models and VRAM usage. | ollama ps |
ollama push |
Push a model. Upload your custom model to ollama.com. | ollama push my-username/my-model |
ollama serve |
Start the service. Start Ollama's API service (usually runs automatically in the background). | ollama serve |
ollama help |
Help. View help information for any command. | ollama help run |
1. Pulling and Deleting Models
pull
Pull a remote model to the local machine.
ollama pull <model>
rm / remove
Delete a local model.
ollama rm <model>
list / ls
List all local models.
ollama list
2. Running Models
run
Run the model in interactive mode without exiting.
ollama run <model>
Can include system information and prompt:
ollama run <model> -s "<system>" -p "<prompt>"
run + script
Read prompt from a file:
ollama run <model> < input.txt
When you enterollama runAfter entering the chat interface, you are no longer operating the command line, but talking to the AI. At this point you can use the/shortcut commands starting with a prefix to control the conversation:
/byeor/exit:Most important!Exit the chat interface and return to the command line./clear: Clear the current context memory (start a new conversation)./show info: View detailed parameter information of the current model./set parameter seed 123: Set a random seed (advanced technique, used to reproduce results)./help: View all available shortcut commands in the chat.
3. Inference Interface (One-time Execution)
generate
Perform a single inference and output text.
ollama generate <model> -p "<prompt>"
4. Creating and Modifying Models
create
Create a local model from a Modelfile.
ollama create <model-name> -f Modelfile
cp
Copy a model to a new name.
ollama cp <src> <dst>
5. Server Related
serve
Start the Ollama local service (default 11434).
ollama serve
run serverless
Whenollama runIt automatically starts the background service; no need to execute it separately.
6. Model Information
show
View model metadata, parameters, and templates.
ollama show <model>
7. Special Parameters
Most of these parameters can be used in run/generate:
--num-predict <number> 限制输出 token 数 --temperature <float> 控制随机性 --top-k <int> 采样范围 --top-p <float> 核采样 --seed <int> 固定随机性 --format json 输出 JSON --keepalive <seconds> 会话保持时间
8. Modelfile Directives
Used when building a model:
- FROM <model>: Base model
- SYSTEM "xxx": Set system prompt
- PARAMETER key=value: Set default parameters
- TEMPLATE "xxx": Customize Chat template
- LICENSE "xxx": Set License
- ADAPTER <file> / WEIGHTS <file>: Load LoRA or additional weights
9. API (when serve is running)
REST endpointshttp://localhost:11434/api:
/api/generate: Text generation/api/chat: Chat streaming interface/api/pull: Remote pull/api/tags: Local model list
Example call (curl):
curl http://localhost:11434/api/generate \
-d '{"model":"qwen2.5","prompt":"hello"}'
10. Advanced
Run with custom parameters:
ollama run <model> --temperature 0.2 --top-p 0.9
Persistent session (preserving context):
Sessions are automatically managed by the model's internal cache, without the need for additional commands.