研究发现AI技能维护主要靠人主导,AI仅辅助,且当前无法可靠衡量规则性。
Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance
- 通过分析5个开源技能库的提交历史,发现所有实质性修改均由人类发起或合并。
- 62%的修改有AI协同痕迹,但多数修改是内容增补与修正,属真实维护行为。
- 现有方法难以准确判断修改是否符合规则,测量标准仍不明确。
长期运行的LLM代理越来越依赖外部技能文件(如SKILL.md)来持续复用能力。这些技能需随工具和使用模式变化而不断修正、扩展和整合。尽管已有研究尝试自动化技能维护,但多基于自动化基准评估,且将人工维护视为未被量化的瓶颈。本文直接研究这一缺失环节:挖掘五个公共AI技能库的完整提交历史,涵盖873次提交、143个技能文件及254次实质性更新(2025年10月至2026年6月)。采用预注册的治理、操作与触发证据编码手册对每次修改进行标注。结果表明:第一,所有实质性修改均来自命名的人类账户,其中62%包含AI协作者标识;第二,这些修改为真实维护行为——抽样验证显示多数改变技能内容,操作类型以新增与修正为主;第三,预注册的规则性轴线未能通过可靠性检验,从提交记录中可靠编码规则性仍是开放问题。我们公开数据集、编码手册、挖掘脚本与重放协议,供未来自动化维护者参考。对于自演化代理而言,当前公开技能维护更像一个由人主导、AI辅助的循环过程,未来维护系统必须在此基础上建立并验证。
原文摘要 · Abstract (English)
Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded, and consolidated as tools and usage patterns shift over deployment. Recent work seeks to automate skill curation, but it largely evaluates against automated baselines and treats human maintenance as an unmeasured bottleneck. We study that missing process directly. We mine the full commit histories of five public AI-skill repositories, a purposive sample of AI-tooling organizations, covering 873 commits, 143 skill files, and 254 substantive post-creation edits from October 2025 to June 2026. We code each edit with pre-registered governance, operation, and trigger-evidence codebooks. Three findings emerge. First, every substantive edit is authored or merged through a named human account, while 62% carry an AI co-author trailer, with large repository-level variation. Second, these edits are genuine curation: an audited sample shows that most change skill content, and the coded operations are dominated by additions and corrections. Third, a pre-registered rule-likeness axis fails its reliability gate; reliably coding rule-likeness from commit artifacts remains an open measurement problem. We release the corpus, codebooks, mining scripts, and a replay protocol for automated skill curators. For self-evolving agents, public skill maintenance currently looks less like an autonomous pipeline than a human-governed, AI-assisted loop that future curators must measure against and operate within.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。