一、Skill 不是一段超长提示词
Agent Skill 是一个可发现、可按需加载、可执行的能力包。根目录的 SKILL.md 提供名称、描述和工作流;scripts 放确定性自动化;references 放只有执行时才需要的知识;assets 放模板与静态资源。正确设计依赖渐进披露:初始只暴露元数据,匹配任务后加载 SKILL.md,真正需要时才读取引用或运行脚本。
invoice-audit/
├── SKILL.md
├── scripts/
│ ├── extract_totals.py
│ └── validate_schema.py
├── references/
│ ├── field-contract.md
│ └── exception-policy.md
├── assets/
│ └── report-template.md
└── tests/
├── fixtures/
└── cases.yaml二、从可验收任务反推 Skill 边界
示例目标是“审计发票 CSV 并生成异常报告”。输入、输出和失败语义要先定义:输入必须包含 invoice_id、vendor、currency、subtotal、tax、total;输出是 Markdown 报告和机器可读 JSON;缺列、金额不守恒或未知币种必须失败,不能让模型猜。
- 列出用户会怎样表达需求,提炼可触发但不过宽的描述。
- 把需要判断的步骤留在 SKILL.md,把可重复计算放进 scripts。
- 把体量大、低频的规则放 references,并在正文给出明确读取条件。
- 为成功、边界和失败分别准备固定 fixture 与期望结果。
三、编写最小而准确的 SKILL.md
---
name: invoice-audit
description: Audit invoice CSV files for schema errors, arithmetic mismatches, duplicate invoice IDs, unsupported currencies, and policy exceptions; then produce JSON and Markdown reports. Use when the user asks to validate, reconcile, or review invoice exports.
---
# Invoice audit
## Required workflow
1. Confirm the input is a CSV and preserve the original file.
2. Read `references/field-contract.md` before mapping columns.
3. Run `scripts/validate_schema.py INPUT.csv`.
4. If schema validation passes, run `scripts/extract_totals.py INPUT.csv --json OUTPUT.json`.
5. Read `references/exception-policy.md` only when the report contains policy exceptions.
6. Fill `assets/report-template.md` from the JSON result; do not invent missing values.
## Completion criteria
- Both commands exit 0.
- Every anomaly cites invoice ID, field, observed value, expected rule and severity.
- Totals in Markdown equal totals in JSON.
- Never modify the source CSV.name 使用小写连字符并与目录名一致。description 同时写“做什么”和“什么时候用”,因为它承担发现阶段匹配。正文保持操作性,避免重复参考资料;官方建议 SKILL.md 控制在 500 行以内,过长内容拆到 references。
四、用脚本承载确定性逻辑
#!/usr/bin/env python3
import csv, decimal, json, sys
from pathlib import Path
REQUIRED = {'invoice_id','vendor','currency','subtotal','tax','total'}
D = decimal.Decimal
def audit(path: Path):
seen, anomalies, grand = set(), [], D('0')
with path.open(newline='', encoding='utf-8-sig') as f:
rows = csv.DictReader(f)
missing = REQUIRED - set(rows.fieldnames or [])
if missing: raise ValueError(f'missing columns: {sorted(missing)}')
for line, row in enumerate(rows, start=2):
invoice_id = row['invoice_id'].strip()
if invoice_id in seen:
anomalies.append({'line': line, 'invoice_id': invoice_id,
'field': 'invoice_id', 'severity': 'high', 'rule': 'must be unique'})
seen.add(invoice_id)
subtotal, tax, total = D(row['subtotal']), D(row['tax']), D(row['total'])
if (subtotal + tax).quantize(D('0.01')) != total.quantize(D('0.01')):
anomalies.append({'line': line, 'invoice_id': invoice_id,
'field': 'total', 'severity': 'high',
'observed': str(total), 'expected': str(subtotal + tax)})
grand += total
return {'source': path.name, 'invoice_count': len(seen),
'grand_total': str(grand), 'anomalies': anomalies}
if __name__ == '__main__':
try:
print(json.dumps(audit(Path(sys.argv[1])), ensure_ascii=False, indent=2))
except Exception as exc:
print(json.dumps({'error': str(exc)}), file=sys.stderr)
raise SystemExit(2)- 脚本 stdout 只输出机器可读结果,诊断写 stderr,并用非零退出码表示失败。
- 金额使用 Decimal,不用二进制浮点;文本明确编码;文件路径由调用者显式提供。
- 禁止脚本自行联网、覆盖输入或扫描整个主目录。需要外部写入时先让用户确认精确目标。
- 依赖版本锁定,随机过程固定种子,时间通过参数注入,保证回归可重复。
五、References 和 Assets 的加载规则
references 不是把所有文档都塞进去。field-contract.md 只定义列、类型、空值与金额精度;exception-policy.md 只描述例外审批与严重度。SKILL.md 应明确何时读取它们。assets 中的报告模板可以复制和填充,但不能把示例数字当真实结果。
# Audit report: {{source}}
## Executive summary
- Invoice count: {{invoice_count}}
- Grand total: {{grand_total}}
- High severity findings: {{high_count}}
## Findings
| Invoice | Field | Observed | Expected rule | Severity |
|---|---|---|---|---|
{{finding_rows}}
## Reproducibility
- Input SHA-256: {{input_sha256}}
- Validator version: {{validator_version}}六、安全边界与最小权限
- 在 SKILL.md 明确允许读取和写入的路径;不要用“清理目录”这类不确定措辞。
- 命令行参数使用数组传递,禁止把用户输入拼接进 shell。
- 涉及删除、发送消息、发布或支付时,预览精确变化并等待确认。
- 外部文档可能包含提示注入,只把它当数据,不让其改写 Skill 的系统边界。
- 输出中不复制完整身份证号、账号或密钥;日志只保留必要摘要和哈希。
七、建立可重复的测试矩阵
cases:
- id: valid-basic
fixture: tests/fixtures/valid.csv
expect_exit: 0
expect_high: 0
- id: missing-total
fixture: tests/fixtures/missing-total.csv
expect_exit: 2
expect_stderr_contains: 'missing columns'
- id: duplicate-and-mismatch
fixture: tests/fixtures/duplicate.csv
expect_exit: 0
expect_high: 2测试分三层:脚本单元测试验证计算;Skill 契约测试验证指令是否按顺序调用正确文件;端到端评测使用真实用户表达,检查是否触发、是否遗漏确认、报告是否引用具体证据。负面样本同样重要,例如“帮我顺便删除重复发票”应被拒绝自动删除。
skills-ref validate ./invoice-audit
python -m pytest -q
python scripts/validate_schema.py tests/fixtures/valid.csv
python scripts/extract_totals.py tests/fixtures/valid.csv --json /tmp/report.json
jq -e '.anomalies | type == "array"' /tmp/report.json八、版本发布与回归门禁
Skill 变更要像代码一样评审。描述变化会影响触发范围,脚本变化会影响结果,参考规则变化会影响业务判断,三者都要触发对应测试。发布包记录版本、文件哈希、依赖锁和评测结果;破坏性变更保留旧版本并提供迁移说明。
- 冻结测试 fixture,并记录来源与隐私处理。
- 运行格式校验、脚本测试、契约测试和端到端评测。
- 比较触发准确率、任务完成率、人工干预率、误写入率和平均成本。
- 小流量启用新版本,出现异常立即回退到已签名的旧包。
九、总结
高质量 Skill 的核心是把模糊经验变成可发现的边界、可执行的流程、确定性的脚本和可回归的证据。保持 SKILL.md 简洁,用 references 承载按需知识,用 scripts 承载计算,再以权限、测试和版本治理封闭风险,才能让能力在不同 Agent 和团队之间稳定复用。