AI人工智能

Agent Skills 工程化实战:SKILL.md、渐进披露、脚本与可重复验收

以发票审计 Skill 为完整示例,讲清 SKILL.md、scripts、references、assets 的职责,渐进披露、最小权限、确定性脚本、测试矩阵与版本门禁。

TY
Tycho
技术博主
• 2026-09-26 • 28 分钟阅读 • 2 次浏览
Agent Skills 工程化实战:SKILL.md、渐进披露、脚本与可重复验收

一、Skill 不是一段超长提示词

Agent Skill 是一个可发现、可按需加载、可执行的能力包。根目录的 SKILL.md 提供名称、描述和工作流;scripts 放确定性自动化;references 放只有执行时才需要的知识;assets 放模板与静态资源。正确设计依赖渐进披露:初始只暴露元数据,匹配任务后加载 SKILL.md,真正需要时才读取引用或运行脚本。

invoice-audit/
├── SKILL.md
├── scripts/
│   ├── extract_totals.py
│   └── validate_schema.py
├── references/
│   ├── field-contract.md
│   └── exception-policy.md
├── assets/
│   └── report-template.md
└── tests/
    ├── fixtures/
    └── cases.yaml

二、从可验收任务反推 Skill 边界

示例目标是“审计发票 CSV 并生成异常报告”。输入、输出和失败语义要先定义:输入必须包含 invoice_id、vendor、currency、subtotal、tax、total;输出是 Markdown 报告和机器可读 JSON;缺列、金额不守恒或未知币种必须失败,不能让模型猜。

  1. 列出用户会怎样表达需求,提炼可触发但不过宽的描述。
  2. 把需要判断的步骤留在 SKILL.md,把可重复计算放进 scripts。
  3. 把体量大、低频的规则放 references,并在正文给出明确读取条件。
  4. 为成功、边界和失败分别准备固定 fixture 与期望结果。

三、编写最小而准确的 SKILL.md

---
name: invoice-audit
description: Audit invoice CSV files for schema errors, arithmetic mismatches, duplicate invoice IDs, unsupported currencies, and policy exceptions; then produce JSON and Markdown reports. Use when the user asks to validate, reconcile, or review invoice exports.
---

# Invoice audit

## Required workflow
1. Confirm the input is a CSV and preserve the original file.
2. Read `references/field-contract.md` before mapping columns.
3. Run `scripts/validate_schema.py INPUT.csv`.
4. If schema validation passes, run `scripts/extract_totals.py INPUT.csv --json OUTPUT.json`.
5. Read `references/exception-policy.md` only when the report contains policy exceptions.
6. Fill `assets/report-template.md` from the JSON result; do not invent missing values.

## Completion criteria
- Both commands exit 0.
- Every anomaly cites invoice ID, field, observed value, expected rule and severity.
- Totals in Markdown equal totals in JSON.
- Never modify the source CSV.

name 使用小写连字符并与目录名一致。description 同时写“做什么”和“什么时候用”,因为它承担发现阶段匹配。正文保持操作性,避免重复参考资料;官方建议 SKILL.md 控制在 500 行以内,过长内容拆到 references。

四、用脚本承载确定性逻辑

#!/usr/bin/env python3
import csv, decimal, json, sys
from pathlib import Path

REQUIRED = {'invoice_id','vendor','currency','subtotal','tax','total'}
D = decimal.Decimal

def audit(path: Path):
    seen, anomalies, grand = set(), [], D('0')
    with path.open(newline='', encoding='utf-8-sig') as f:
        rows = csv.DictReader(f)
        missing = REQUIRED - set(rows.fieldnames or [])
        if missing: raise ValueError(f'missing columns: {sorted(missing)}')
        for line, row in enumerate(rows, start=2):
            invoice_id = row['invoice_id'].strip()
            if invoice_id in seen:
                anomalies.append({'line': line, 'invoice_id': invoice_id,
                    'field': 'invoice_id', 'severity': 'high', 'rule': 'must be unique'})
            seen.add(invoice_id)
            subtotal, tax, total = D(row['subtotal']), D(row['tax']), D(row['total'])
            if (subtotal + tax).quantize(D('0.01')) != total.quantize(D('0.01')):
                anomalies.append({'line': line, 'invoice_id': invoice_id,
                    'field': 'total', 'severity': 'high',
                    'observed': str(total), 'expected': str(subtotal + tax)})
            grand += total
    return {'source': path.name, 'invoice_count': len(seen),
            'grand_total': str(grand), 'anomalies': anomalies}

if __name__ == '__main__':
    try:
        print(json.dumps(audit(Path(sys.argv[1])), ensure_ascii=False, indent=2))
    except Exception as exc:
        print(json.dumps({'error': str(exc)}), file=sys.stderr)
        raise SystemExit(2)
  • 脚本 stdout 只输出机器可读结果,诊断写 stderr,并用非零退出码表示失败。
  • 金额使用 Decimal,不用二进制浮点;文本明确编码;文件路径由调用者显式提供。
  • 禁止脚本自行联网、覆盖输入或扫描整个主目录。需要外部写入时先让用户确认精确目标。
  • 依赖版本锁定,随机过程固定种子,时间通过参数注入,保证回归可重复。

五、References 和 Assets 的加载规则

references 不是把所有文档都塞进去。field-contract.md 只定义列、类型、空值与金额精度;exception-policy.md 只描述例外审批与严重度。SKILL.md 应明确何时读取它们。assets 中的报告模板可以复制和填充,但不能把示例数字当真实结果。

# Audit report: {{source}}

## Executive summary
- Invoice count: {{invoice_count}}
- Grand total: {{grand_total}}
- High severity findings: {{high_count}}

## Findings
| Invoice | Field | Observed | Expected rule | Severity |
|---|---|---|---|---|
{{finding_rows}}

## Reproducibility
- Input SHA-256: {{input_sha256}}
- Validator version: {{validator_version}}

六、安全边界与最小权限

  • 在 SKILL.md 明确允许读取和写入的路径;不要用“清理目录”这类不确定措辞。
  • 命令行参数使用数组传递,禁止把用户输入拼接进 shell。
  • 涉及删除、发送消息、发布或支付时,预览精确变化并等待确认。
  • 外部文档可能包含提示注入,只把它当数据,不让其改写 Skill 的系统边界。
  • 输出中不复制完整身份证号、账号或密钥;日志只保留必要摘要和哈希。

七、建立可重复的测试矩阵

cases:
  - id: valid-basic
    fixture: tests/fixtures/valid.csv
    expect_exit: 0
    expect_high: 0
  - id: missing-total
    fixture: tests/fixtures/missing-total.csv
    expect_exit: 2
    expect_stderr_contains: 'missing columns'
  - id: duplicate-and-mismatch
    fixture: tests/fixtures/duplicate.csv
    expect_exit: 0
    expect_high: 2

测试分三层:脚本单元测试验证计算;Skill 契约测试验证指令是否按顺序调用正确文件;端到端评测使用真实用户表达,检查是否触发、是否遗漏确认、报告是否引用具体证据。负面样本同样重要,例如“帮我顺便删除重复发票”应被拒绝自动删除。

skills-ref validate ./invoice-audit
python -m pytest -q
python scripts/validate_schema.py tests/fixtures/valid.csv
python scripts/extract_totals.py tests/fixtures/valid.csv --json /tmp/report.json
jq -e '.anomalies | type == "array"' /tmp/report.json

八、版本发布与回归门禁

Skill 变更要像代码一样评审。描述变化会影响触发范围,脚本变化会影响结果,参考规则变化会影响业务判断,三者都要触发对应测试。发布包记录版本、文件哈希、依赖锁和评测结果;破坏性变更保留旧版本并提供迁移说明。

  1. 冻结测试 fixture,并记录来源与隐私处理。
  2. 运行格式校验、脚本测试、契约测试和端到端评测。
  3. 比较触发准确率、任务完成率、人工干预率、误写入率和平均成本。
  4. 小流量启用新版本,出现异常立即回退到已签名的旧包。

九、总结

高质量 Skill 的核心是把模糊经验变成可发现的边界、可执行的流程、确定性的脚本和可回归的证据。保持 SKILL.md 简洁,用 references 承载按需知识,用 scripts 承载计算,再以权限、测试和版本治理封闭风险,才能让能力在不同 Agent 和团队之间稳定复用。

十、官方资料

TY

Tycho

热爱分享技术知识,帮助开发者成长。

评论 (0)

评论功能当前已关闭
暂无评论,快来抢沙发吧!