优化大模型智能体技能的呈现方式,提升任务成功率
SkillAlign: Aligning Skill Interfaces for LLM-based Agents

- 将技能视为多视角流程卡片,动态调整展示形式
- 压缩展示可比完整注入更高效,最高提升17%成功率
- 适合研究智能体技能接口设计与自适应交互的学者
语言模型智能体依赖技能——可复用的过程性知识,用于推理、工具使用和交互。现有研究关注技能的获取、检索、压缩与组合,但通常假设一旦选定技能,其与智能体的接口即固定不变。我们指出,这一假设忽略了技能效用的关键来源:同一技能因呈现方式不同,可能助益、干扰或误导智能体。为此,提出SkillAlign框架,将候选技能表示为多视角流程卡片,并通过全指令、提示、压缩摘要、工作流或无暴露等不同方式呈现。该方法支持反事实评估,在任务、智能体和候选技能固定时,仅改变呈现接口。在ALFWorld和SkillsBench上验证显示,呈现形式显著影响任务成功率与上下文开销;紧凑的top-k呈现甚至优于完整库注入。进一步在ALFWorld上进行基于回放的策略学习分析表明,自适应呈现包含可学习信号,但远未达到理想选择水平。结果表明,技能增强型智能体不仅需选择合适技能,还需优化其呈现方式。
原文摘要 · Abstract (English)
Language-model agents increasingly rely on skills: reusable procedural knowledge for reasoning, tool use, and interaction. Existing work studies how skills are acquired, retrieved, compressed, or composed, but often assumes that once a skill is selected, its interface to the agent is fixed. We argue that this overlooks a key source of skill utility: the same skill can help, distract, or mislead depending on how it is exposed. We propose SkillAlign, a provider-agnostic framework that represents candidate skills as multi-view procedural cards and renders them through alternative exposure interfaces, including full instructions, hints, compressed summaries, workflows, or no exposure. This enables counterfactual evaluation where the task, agent, and candidate skills are fixed while only the exposure interface varies. Across ALFWorld and SkillsBench, we show that exposure form substantially affects task success and rendered context cost, and that compact top-k exposure can outperform full-library injection. We further conduct a replay-based policy-learning analysis on ALFWorld, showing that adaptive exposure contains learnable signal but remains far from oracle selection. Our results suggest that skill-augmented agents should optimize not only which skills to use, but also how those skills are presented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。