Skills Unit Testing
After writing a Skill, how can you confirm it actually works as expected? Manual testing is time-consuming and error-prone.
This article introduces how to build systematic test cases for a Skill, from script unit testing to end-to-end trigger testing.
Two Levels of Skill Testing
| Level | Test Target | Method |
|---|---|---|
| Script Unit Testing | Python/JS scripts under scripts/ | pytest / standard unittest |
| Trigger-based End-to-End Testing | Whether the Skill is triggered correctly and whether the output meets expectations | skill-creator's eval framework |
Level 1: Script Unit Testing
Scripts are the most deterministic part of a Skill and are well suited to coverage with a standard unit testing framework.
Example
import pytest
import os
import tempfile
import csv
# Module under test
import sys
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from clean_data import remove_duplicates, fill_nulls, strip_whitespace
# ── Test: Deduplication ──────────────────────────────────────────
def test_remove_duplicates_basic():
"""Normal case: duplicate rows should be removed"""
rows = [
{"name": "example", "score": "90"},
{"name": "example", "score": "90"}, # Duplicate row
{"name": "EXAMPLE", "score": "80"},
]
result = remove_duplicates(rows)
assert len(result) == 2, "1 duplicate row should be removed"
def test_remove_duplicates_empty():
"""Edge case: empty list should not raise an error"""
result = remove_duplicates([])
assert result == []
# ── Test: Null Value Filling ──────────────────────────────────────
def test_fill_nulls_numeric():
"""Null values in numeric columns should be filled with 0"""
rows = [{"value": "10"}, {"value": ""}, {"value": "20"}]
result = fill_nulls(rows, col="value", fill_with="0")
assert result[1]["value"] == "0"
# ── Test: Trimming Spaces ──────────────────────────────────────
def test_strip_whitespace():
"""Leading and trailing spaces in strings should be removed"""
rows = [{"name": " example "}, {"name": "EXAMPLE"}]
result = strip_whitespace(rows, col="name")
assert result[0]["name"] == "example"
assert result[1]["name"] == "EXAMPLE"
# ── Integration Test: Reading a Real File ──────────────────────────────
def test_process_real_file():
"""Create a temporary CSV file and test the complete processing flow"""
with tempfile.NamedTemporaryFile(mode="w", suffix=".csv",
delete=False, newline="") as f:
writer = csv.DictWriter(f, fieldnames=["name", "score"])
writer.writeheader()
writer.writerow({"name": " example ", "score": "90"})
writer.writerow({"name": " example ", "score": "90"}) # Duplicate
writer.writerow({"name": "EXAMPLE", "score": ""}) # Null value
tmp_path = f.name
try:
from clean_data import process_file
result = process_file(tmp_path)
assert result["removed_rows"] == 1
assert result["fixed_values"] == 1
finally:
os.unlink(tmp_path)
Run the tests:
# 进入 Skill 目录后运行所有测试 cd my-skill/ pytest scripts/tests/ -v # 生成覆盖率报告 pytest scripts/tests/ --cov=scripts --cov-report=term-missing
Output:
scripts/tests/test_clean_data.py::test_remove_duplicates_basic PASSED scripts/tests/test_clean_data.py::test_remove_duplicates_empty PASSED scripts/tests/test_clean_data.py::test_fill_nulls_numeric PASSED scripts/tests/test_clean_data.py::test_strip_whitespace PASSED scripts/tests/test_clean_data.py::test_process_real_file PASSED 5 passed in 0.12s
Level 2: Trigger-based End-to-End Testing
Trigger testing verifies whether Claude will use the Skill when a user makes a request.
Test cases are defined in JSON format, containing user prompts and expected behavior assertions.
Example
{
"id": "trigger_basic",
"prompt": "Help me analyze this sales data CSV, find monthly trends and generate a statistical summary",
"assertions": [
{
"type": "skill_triggered",
"skill": "csv-analyzer",
"description": "Complex analysis requests should trigger csv-analyzer"
}
]
},
{
"id": "trigger_with_file",
"prompt": "I uploaded example_sales.csv, please analyze the data distribution of each column",
"assertions": [
{
"type": "skill_triggered",
"skill": "csv-analyzer"
},
{
"type": "output_contains",
"keyword": "Statistics",
"description": "The output should contain statistical information"
}
]
},
{
"id": "no_trigger_simple",
"prompt": "What format is CSV?",
"assertions": [
{
"type": "skill_not_triggered",
"skill": "csv-analyzer",
"description": "Simple knowledge questions should not trigger the Skill"
}
]
}
]
Run trigger tests (requires the eval tool provided by skill-creator):
Example
python -m scripts.run_eval \
--eval-set evals/trigger-eval.json \
--skill-path csv-analyzer/ \
--model claude-sonnet-4-20250514
# Generate an HTML report for manual review
python eval-viewer/generate_review.py \
--results evals/results/ \
--output evals/review.html
Test case prompts should be sufficiently complex. For simple requests like "read CSV", even if triggering fails, it doesn't mean the Skill has a problem, because Claude can handle it on its own. Good test cases should be multi-step requests users would actually make, with clear output requirements.
Test Case Coverage
| Type | Scenarios to cover |
|---|---|
| Positive triggering | Complex tasks, containing keywords, clear output format requirements |
| Negative triggering | Simple Q&A, only mentioning relevant terms but not needing the Skill |
| Boundary inputs | Empty files, oversized files, files with unsupported formats |
| Error recovery | Whether the error message is clear when the file path is wrong |