arXiv:2605.22634cs.SEcs.AI2026-05

为企业AI代理设计可读的任务契约,提升指令清晰度与执行质量。

Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

  • 用契约化方式组织技能文件,明确任务目标与边界
  • 实测使输出质量均值从4.692升至4.914,关键错误率降至0.013
  • 适合需规范协作的大型企业AI系统开发者

技能已成为智能体指令、工作流、脚本和参考材料的实用封装方式。在企业场景中,技能不仅需提供任务指导,还需明示目标、输入范围、权限、人工审批点、证据要求、输出契约、质量标准、验证步骤及交接规则。本文提出‘契约技能’(Contractual Skills),一种受GovernSpec启发的设计框架,将SKILL.md文件组织为可读的任务契约,同时保持轻量级技能发现与渐进加载能力。该框架厘清了契约技能、GovernSpec YAML契约、模型上下文协议(MCP)接口、工具适配器、运行时护栏、追踪与评估系统之间的边界。通过三项离线实证研究验证:第一项文本生成实验涵盖三个企业技能、十五个合成任务、四种指令条件、八种生成模型,产出960条输出与1680条交叉评分记录;第二项公开技能A/B对比实验中,八个公开技能经契约重写后,在四十八个合成任务、六种模型、两次重复下产生1152条输出与两份完整评分文件,契约技能使平均质量从4.692提升至4.914,关键错误率由0.083降至0.013;第三项离线工具调用挑战包含八种模型与192条模拟调用记录。结果表明,契约技能更应被视为一种治理层,明确任务意图、边界与验收标准,而非独立安全机制。

原文摘要 · Abstract (English)

Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings, however, a skill often needs to express more than task guidance: goals, input boundaries, permissions, human approval points, evidence requirements, output contracts, quality criteria, verification steps, and handoff rules. This paper proposes contractual skills, a GovernSpec-inspired design framework for organizing SKILL.md files as readable task contracts while preserving lightweight skill discovery and progressive loading. The framework clarifies the boundary between contractual skills, GovernSpec YAML contracts, Model Context Protocol (MCP) surfaces, tool adapters, runtime guardrails, tracing, and evaluation systems. We evaluate the framework with three offline empirical studies. The first text-generation experiment covers three enterprise skills, fifteen synthetic tasks, four instruction conditions, and eight generation models, producing 960 outputs and 1680 cross-judge score records. The second study is a public-skill A/B expansion: eight public skills are compared with contractual rewrites across forty-eight synthetic tasks, six generation models, two repeats, 1152 outputs, and two complete judge files. In this setting, contractual skills raise mean quality from 4.692 to 4.914 and reduce critical-error rate from 0.083 to 0.013. The third study is an offline tool-calling challenge with eight models and 192 simulated tool-call records. The results suggest that contractual skills are best understood as a governance layer that makes task intent, boundaries, and acceptance criteria explicit, not as a standalone safety mechanism.

企业AI任务契约治理框架技能管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。