Skills Description Optimization Tips

The description field is the sole basis for determining whether a Skill is triggered.

A well-written description allows the Skill to be triggered at the right time and avoids interfering with Claude in irrelevant scenarios.


The Role and Limitations of Description

When deciding whether to use a Skill, Claude only reads the description field and does not read the SKILL.md content in advance.

This means the description needs to accomplish two things at once: explain what the Skill can do, and indicate when it should be used.

The ideal length for a description is 50-150 characters. Too short leads to inaccurate triggering; too long may exceed context limits, and Claude will find it difficult to quickly scan and compare.


The Four-Element Structure of Description

ElementPurposeExample
One-sentence core functionMost directly states what the Skill is"Process spreadsheet data in Excel and CSV formats"
List of specific capabilitiesList 3-5 typical operations"Including data cleaning, statistical analysis, chart generation"
Trigger scenario descriptionTells Claude when it should be used"Trigger when the user mentions data analysis, reports, or statistics"
Promotional statement (optional)Solves the "low trigger rate" problem"Even if the user does not explicitly mention the format, it should be used whenever tabular data is involved"
---
name: data-analyzer
description: >
  处理 Excel(.xlsx)和 CSV 格式的表格数据,包括数据读取、
  清洗去重、统计分析(均值/中位数/分布)、生成可视化报告。
  当用户需要分析数据文件、生成统计报告、处理表格、
  查找数据规律时应优先使用此 Skill,即使用户只是上传了
  一个数据文件并询问"帮我看看这个"。
---

Common Description Writing Problems

ProblemExample (incorrect)Correction direction
Only writing functions, not trigger scenarios"A tool for analyzing data"Add "trigger when the user..."
Description is too broad"A general-purpose tool for processing files"List specific file types and operations
Using Skill internal implementation terminology"Call pandas for analysis"Change to a business description from the user's perspective
Unclear boundaries with other SkillsBoth Skills say "process documents"Clearly distinguish each one's file types or operation types

Optimizing Description with Automated Tools

skill-creator provides a description auto-optimization script that uses iterative testing to find the writing style with the highest trigger rate.

Step 1: Prepare Test Cases

Example

[
  {
    "id": "should_trigger_01",
    "prompt": "Help me analyze example_sales.csv and find the monthly sales trend",
    "expected": true,
    "note": Typical data analysis request
  },
  {
    "id": "should_trigger_02",
    "prompt": "Which columns in this Excel file have empty values? Help me count them.",
    "expected": true,
    "note": Involves data quality checking
  },
  {
    "id": "should_not_trigger_01",
    "prompt": "What is the difference between Excel and CSV?",
    "expected": false,
    "note": Knowledge Q&A, no Skill needed
  },
  {
    "id": "should_not_trigger_02",
    "prompt": "Help me write an email",
    "expected": false,
    "note": Completely unrelated task
  }
]

Step 2: Run the Optimization Loop

Example

# Run description auto-optimization
# The script will automatically: evaluate → modify → re-evaluate, looping 5 times
python -m scripts.run_loop \
  --eval-set evals/trigger-eval.json \
  --skill-path data-analyzer/ \
  --model claude-sonnet-4-20250514 \
  --max-iterations 5 \
  --verbose
迭代 1:训练集得分 0.62,测试集得分 0.60
迭代 2:训练集得分 0.75,测试集得分 0.72
迭代 3:训练集得分 0.87,测试集得分 0.85
迭代 4:训练集得分 0.90,测试集得分 0.88
迭代 5:训练集得分 0.91,测试集得分 0.89

最优 description(测试集得分 0.89)已保存。

Step 3: Apply the Optimal Result

Example

# The script outputs the best_description field; replace it into the description in SKILL.md
# Before optimization
grep "description:" data-analyzer/SKILL.md

# Manually or with a script, replace it with the content of best_description
# Repackage the Skill after replacement
python -m scripts.package_skill data-analyzer/

A/B Comparison Testing

When unsure which writing style is better, you can perform a blind test comparison of the two versions of the description.

Example

# Version A: current description
# Version B: candidate new description

# Run evaluations separately
python -m scripts.run_eval \
  --eval-set evals/trigger-eval.json \
  --skill-path data-analyzer-v1/ \
  --output evals/results_v1.json

python -m scripts.run_eval \
  --eval-set evals/trigger-eval.json \
  --skill-path data-analyzer-v2/ \
  --output evals/results_v2.json

# Generate a comparison report
python eval-viewer/generate_review.py \
  --results evals/results_v1.json evals/results_v2.json \
  --compare \
  --output evals/comparison.html

Evaluation should be based on the test set score, not the training set score. If the training set score is high but the test set score is low, it means the description has overfit the test cases, and its performance in real scenarios may deteriorate.

Other Extensions