用自演化评分体系精准评估智能体用技能的全过程
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

- 从真实运行中提取四维评分标准:选技、执行、组合、反思
- 发现最终成功无法揭示的技能误用问题,提升评估精度
- 适合训练和优化大模型智能体的技能使用能力
技能正成为大模型智能体可复用的操作层,涵盖标准流程、领域规则、工具流程、脚本与验证逻辑。在真实技能库中,技能重叠导致可靠使用困难。仅依赖最终验证结果进行评估与训练过于粗糙——智能体可能通过试错选择干扰项、跳过必要步骤、错误组合流程或遗漏最终检查而侥幸成功。本文提出SkillCoach,一种自演化评分框架,用于评估与增强智能体的技能使用能力。该框架从真实运行轨迹中提取基于技能的流程评分标准,从四个维度评估轨迹:技能选择、技能执行、技能组合及技能驱动的反思。保持外部验证器作为独立结果信号,使过程质量与偶然任务成功相区分。演化的评分标准进一步作为过程监督信号,筛选高质量训练轨迹。实验表明,演化后的评分标准显著提升评估质量,暴露了最终准确率无法察觉的失败模式,并为增强技能使用提供了比仅依赖结果的过滤更强的监督信号。
原文摘要 · Abstract (English)
Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use difficult. Final verifier success is too coarse for both evaluation and training, since an agent may pass through trial and error while selecting distractor skills, skipping required steps, composing workflows incorrectly or omitting final checks. We introduce SkillCoach, a self-evolving rubric framework for evaluating and enhancing agentic skill-use. SkillCoach derives skill-grounded process rubrics from real rollouts and evaluates trajectories along four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection. It keeps the external verifier as a separate outcome signal, allowing process quality to be distinguished from accidental task success. The evolved rubrics further serve as process supervision for selecting high-quality training trajectories. Experiments show that evolved rubrics substantially improve evaluation quality, expose failures hidden by final accuracy, and provide stronger supervision signals than outcome-only filtering for enhancing agentic skill-use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。