实验证明:给大模型提供技能文档能显著提升任务成功率,但呈现方式影响不大。
Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study
- 用30个任务测试不同技能呈现方式对模型表现的影响
- 有技能时任务通过率平均提升18%~36%,效果明显
- 技能粒度细节变化影响小且不一致,适合做系统评估的基准研究
技能文档在推理时为大语言模型代理提供程序性知识。本文通过固定版本的SkillsBench(一个经官方验证的30任务领域均衡子集),结合两种具备推理能力的模型配置、六种技能条件和每组任务-条件-模型组合5次试验,研究技能呈现粒度对下游任务成功率的影响。结果显示,相比无技能情况,有技能条件下GPT-5.5的任务平均通过率提升26.7至36.0个百分点,DeepSeek V4-Flash提升18.0至26.0个百分点。最终数据共1800行,每模型900行。任务为推理单元,每组内5次试验聚合后,在30个任务上进行配对对比。主要呈现方式对比效果较小且不确定:低抽象引导相较于高抽象引导,GPT-5.5仅提升0.7个百分点,DeepSeek V4-Flash反而下降6.7个百分点,两者95%置信区间均包含零。在中等抽象引导中加入一个示例,相较无示例版本分别提升0.7和1.3个百分点。均值奖励鲁棒性检验结果一致。在该受控子集中,技能可用性显著提升成功率,而呈现粒度的变化则产生微小、不确定且依赖模型的效果。
原文摘要 · Abstract (English)
Skill documents provide procedural knowledge to large-language-model agents at inference time. This article studies whether the presentation granularity of controlled skill knowledge changes downstream task success. The experiment uses a pinned SkillsBench version, a 30-task domain-balanced subset validated by official oracle runs, two reasoning-enabled model configurations, six skill conditions, and five trials per task-condition-model cell. Skill availability is the clearest empirical signal. Relative to no skill, skill conditions increase task-mean pass rate by 26.7 to 36.0 percentage points for GPT-5.5 and by 18.0 to 26.0 percentage points for DeepSeek V4-Flash. The final data contain 1,800 rows, with 900 rows for each model. The task is the inference unit. Five trials are aggregated within each task-condition-model cell before paired contrasts are estimated over 30 tasks. The primary presentation contrasts are smaller and uncertain. Low-abstraction guidance differs from high-abstraction guidance by +0.7 percentage points for GPT-5.5 and -6.7 percentage points for DeepSeek V4-Flash, with both 95% bootstrap confidence intervals crossing zero. Adding one worked example to medium-abstraction guidance differs from the no-example variant by +0.7 and +1.3 percentage points. Mean-reward robustness checks preserve the same substantive conclusion. In this controlled subset, skill availability is associated with higher success than no skill, while the tested presentation-granularity changes yield small, uncertain, and model-dependent effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。