Hermes Agent Toolkit
Tools are the only channel through which Hermes Agent interacts with the external world—reading and writing files, executing commands, searching the web, generating images; all capabilities are exposed through tools.
Hermes has 70+ built-in tools, organized into 28 tool collections. This chapter will analyze the core tools, parameters, and best practices of each tool collection one by one.
Tool system architecture
Before diving into specific tools, let's first understand how tools are organized.
Hermes uses a three-layer structure to manage tools:
工具(Tool) ← 单个可调用函数,如 read_file、web_search
└── 工具集(Toolset) ← 相关工具的逻辑分组,如 file、web
└── 平台配置 ← 哪个平台加载哪些工具集
Key Design: Disabling a tool collection makes its tools completely disappear from the system prompt—not just uncallable, but the Agent is entirely unaware of their existence. This both saves Tokens and prevents the Agent from misusing tools it shouldn't use.
Ways to manage tool sets:
# TUI 交互式管理(推荐) hermes tools # 命令行管理 hermes tools list # 列出所有工具及启用状态 hermes tools enable browser # 启用指定工具集 hermes tools disable audio # 禁用指定工具集
Configure in config.yaml:
# 文件路径:~/.hermes/config.yaml
# 按平台指定启用的工具集
platform_toolsets:
cli:
- hermes-cli # CLI 核心工具集
- file
- terminal
- web
- memory
- skills
Desktop version, in the left-side Skills & Tools menu:

Tools can also be managed on the dashboard, in the left-side skills menu:

File and system tool sets
This type of tool collection enables the Agent to manipulate the local file system, execute commands, and manage processes.
file — file operations
The most fundamental tool collection; the Agent relies on it for nearly all tasks to read and write files.
| Tools | Features | Key Parameters |
|---|---|---|
| read_file | Read file content | file_path (required), offset (starting line), limit (maximum number of lines) |
| write_file | Create or overwrite file | file_path (required), content (required) |
| edit_file | Precisely replace text in files | file_path、old_string、new_string、replace_all |
| list_dir | List directory contents | path (required), recursive (whether to recurse) |
| search_files | Search by file name pattern | pattern (required, supports glob), path (search root directory) |
| search_content | Search text in file content | pattern (required, supports regex), path, file_types |
How edit_file works: It uses exact substring matching to locate the modification point. The Agent provides old_string (text that actually exists in the file), the tool finds it and replaces it with new_string. If old_string matches multiple locations, set replace_all: true to replace all occurrences.
edit_file is the Agent's most commonly used file modification method—it is safer than write_file because if old_string is not found, the operation fails with an error instead of accidentally overwriting the file.
terminal — Shell command execution
Allows the Agent to execute Shell commands in a specified backend environment.
| Tools | Features | Key Parameters |
|---|---|---|
| run_command | Execute a single Shell command | command (required), workdir (working directory), timeout (timeout in milliseconds) |
| run_script | Execute multi-line Shell scripts | script (required, multi-line script content), workdir, timeout |
Commands execute in the configured terminal backend—local, docker (container), ssh (remote), etc.
Execution results include stdout, stderr, and exit code. The Agent can use this information to determine whether the command succeeded.
code — sandbox code execution
Executes code snippets in an isolated sandbox, supporting Python, JavaScript, and Bash.
| Tools | Features | Key Parameters |
|---|---|---|
| execute_code | Run code in sandbox | language (python/js/bash), code (required), timeout |
Difference from terminal: the code tool runs in an independent sandbox and cannot access the file system (unless explicitly mounted). It is suitable for running code that does not require file interaction, such as algorithm validation and data processing.
computer — desktop GUI automation
Controls the mouse, keyboard, and screenshots of the desktop environment.
| Tools | Features |
|---|---|
| screenshot | Capture the screen or a specified region |
| click | Click the mouse at specified coordinates |
| type | Simulate keyboard text input |
| scroll | Scroll mouse wheel |
| move | Move the mouse to specified coordinates |
The computer tool collection is mainly used for GUI application testing and automated operations. If your use case doesn't involve desktop applications, you can disable it to save Tokens.
Network and content tool sets
These tool sets enable the Agent to search the internet, extract web page content, and automate browser operations.
web — Web search and content extraction
The main channel through which the Agent obtains external information.
| Tools | Features | Key Parameters |
|---|---|---|
| web_search | Search internet content | query (required), num_results (number of results) |
| web_extract | Extract the main content of a specified URL | url (required), extract_mode (auto/text/markdown) |
web_search returns a list of search results (title + URL + summary). The Agent typically searches first, then calls web_extract on pages of interest to get detailed content.
A search API must be configured to use it. Multiple providers are supported:
Example
# Configure web search backend
web:
search_provider: firecrawl # or brave, serpapi, tavily
# Set the corresponding API Key in .env
# FIRECRAWL_API_KEY=fc-...
# Or BRAVE_API_KEY=BS...
browser — Browser automation
Full browser control capabilities—navigation, clicking, form filling, screenshots—operating web pages like a human.
| Tools | Features | Key Parameters |
|---|---|---|
| navigate | Navigate to a specified URL | url (required) |
| click | Click page element | selector (CSS selector or text description) |
| fill_form | Fill form fields | fields (mapping from field names to values) |
| screenshot | Capture current page | full_page (whether to take a full-page screenshot) |
| extract | Extract structured data from a page | selector、format(json/text) |
| execute_js | Executes JavaScript in the page | code (required) |
The browser tool collection requires Playwright or a cloud browser backend.
The browser tool set consumes additional resources when launching browser instances. If you only need to extract web page text content, prefer web_extract over browser extract—the former is a lightweight HTTP request, while the latter requires launching a full browser.
vision — image analysis
Allows the Agent to "understand" image content.
| Tools | Features | Key Parameters |
|---|---|---|
| analyze_image | Analyze image content and generate a description | image_path or image_url (required), question (optional, specific question about the image) |
| describe_image | Generate a detailed textual description of an image | image_path or image_url (required) |
The vision tool sends images to a vision-capable LLM (such as claude-sonnet-4-6, gpt-4o), and the model recognizes and describes the image content.
audio — audio processing
Speech-to-text and text-to-speech.
| Tools | Features | Key Parameters |
|---|---|---|
| transcribe | Convert audio files to text | audio_path (required), language (optional) |
| tts | Convert text to speech files | text (required), voice (voice selection), provider |
Agent capability tool sets
This type of tool collection manages the Agent's own state—memory, skills, and sub-Agent collaboration.
memory — memory management
Read and write persistent memory and search cross-session history.
| Tools | Features | Key Parameters |
|---|---|---|
| memory | Manage persistent memory (add, delete, modify) | action(add/replace/remove)、target(memory/user)、content |
| session_search | Full-text search historical sessions | query (required, supports FTS5 syntax) |
The memory tool has no read operation—persistent memory is already injected into the system prompt at session startup, so the Agent doesn't need to read it separately.
skills — skill management
Loads and manages reusable workflow skills.
| Tools | Features | Key Parameters |
|---|---|---|
| skills_list | List all installed skills and their descriptions | None (automatically called at session startup) |
| skill_view | Load the full content of a skill | name (required), path (optional, load references sub-files) |
| skill_manage | Create, modify, or delete skill files | action、name、content |
delegation — sub-agent delegation
Spawns sub-Agents to execute tasks in parallel.
| Tools | Features | Key Parameters |
|---|---|---|
| spawn_agent | Creates a single sub-Agent to execute a task | prompt (required), model (optional) |
| spawn_parallel | Create multiple sub-Agents in parallel | tasks (required, task list) |
Sub-Agents have independent context windows and do not pollute the main Agent's conversation history. After execution, they return a summary of results.
kanban — Multi-agent kanban board
A structured multi-Agent collaboration system that requires explicit enabling (all/* does not automatically turn it on).
| Tools | Features |
|---|---|
| task_create | Create a new task in the kanban board |
| task_update | Update task status (todo/in_progress/done) |
| task_list | List all tasks of the tenant |
| task_assign | Assigns tasks to designated worker Agents |
Kanban is the core mechanism for multi-Profile collaboration—one orchestrator Profile creates tasks, and multiple worker Profiles claim and execute them. This requires the gateways of multiple Profiles to run simultaneously.
Media creation toolset
This type of tool collection typically requires a Nous Portal subscription or the corresponding third-party API Key.
image_gen — AI image generation
Generates images from text descriptions, supporting 9 models.
| Tools | Features | Key Parameters |
|---|---|---|
| generate_image | Generate an image based on a Prompt | prompt(Required)、model(flux/gpt-image/ideogram etc.)、size、style |
tts — text-to-speech
TTS functionality independent of the audio tool collection, offering a richer selection of providers.
| Tools | Features |
|---|---|
| text_to_speech | Convert text to speech files |
| list_voices | List available voice timbres |
Custom toolset
You can combine existing tool collections into project-specific custom tool collections:
Example
# Define project-specific toolset combination
custom_toolsets:
# Tools needed for data science projects
data-science:
- file
- terminal
- code
- web
# Pure writing projects only need file operations
writing:
- file
- web
- memory
# Ops projects need terminal + files, not a browser
devops:
- file
- terminal
- memory
When using, through--toolsetsParameter specification:
Example
hermes --toolsets data-science chat
# Or set as default in Profile configuration
coder config set toolsets data-science
Tool call flow
Understanding the complete flow of a tool call helps to understand the Agent's behavior and troubleshoot issues.
A complete tool call goes through the following steps:
- The LLM decides to call a tool while generating its reply, outputting the tool name and parameters
- tool_request Middleware execution (can modify parameters)
- Approval check—if the tool is a dangerous operation, decides whether to show a confirmation prompt based on approvals.mode
- Actual tool execution
- tool_execution Middleware executes (can modify results)
- post_tool_call Hook records telemetry data
- The execution result is returned to the LLM, and the LLM continues generating a reply
If your Middleware modifies tool parameters, the modified parameters are the input for approval and actual execution. If you short-circuit execution in the Middleware (return a custom result), the actual tool will not run.
Tool usage statistics
Knowing which tools are frequently used helps optimize tool collection configuration and Token consumption.
In a session, view the tool call statistics for the current session:
Example
/usage
# Output example:
# Token usage: Input 12,340 | Output 5,678 | Total 18,018
# Tool calls: read_file (15 times) | edit_file (8 times)
# terminal (3 times) | web_search (2 times)
In batch processing scenarios, the statistics.json file records more detailed statistics:
Example
{
"total_tool_calls": 1523,
"tools": {
"read_file": {"count": 520, "success": 510, "failure": 10},
"edit_file": {"count": 380, "success": 375, "failure": 5},
"terminal": {"count": 290, "success": 270, "failure": 20},
"web_search": {"count": 180, "success": 178, "failure": 2}
},
"avg_tool_calls_per_turn": 2.4,
"most_common_tool_sequence": [
"read_file", "edit_file", "terminal"
]
}