arXiv:2603.00718cs.CLcs.SE2026-03被引 27

测试大模型能否像人一样学会组合使用工具并复用技能。

SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?

  • 设计可扩展的工具组合任务,让模型学习抽象和复用高级技能。
  • 使用技能复用后,模型耗能最多降低80%,效率显著提升。
  • 适合研究长期任务规划与智能体自主学习的学者参考。

现实世界中的工具使用智能体需处理长周期、结构重复且需求多样的任务,有效行为不仅依赖原子工具调用,还需抽象和复用高层次工具组合。然而,现有基准大多仅衡量静态工具集下的单次成功率,难以反映智能体习得可复用技能的能力。为此,我们提出SkillCraft基准,专门测试智能体形成与复用高层次工具组合(称为技能)的能力。该基准包含真实、高度组合化的工具使用场景,难度在数量和结构维度上均可调节,旨在激发技能抽象与跨任务复用。我们进一步提出轻量级评估协议,使智能体能自动将原子工具组合为可执行技能,缓存并跨任务复用,从而提升效率并积累持久可复用的技能库。在SkillCraft上评估当前最优智能体,发现使用技能缓存与复用后,令牌消耗最高减少80%。同时,测试时的成功率与工具组合能力强相关,表明组合技能习得是核心能力。

原文摘要 · Abstract (English)

Real-world tool-using agents operate over long-horizon workflows with recurring structure and diverse demands, where effective behavior requires not only invoking atomic tools but also abstracting, and reusing higher-level tool compositions. However, existing benchmarks mainly measure instance-level success under static tool sets, offering limited insight into agents' ability to acquire such reusable skills. We address this gap by introducing SkillCraft, a benchmark explicitly stress-test agent ability to form and reuse higher-level tool compositions, where we call Skills. SkillCraft features realistic, highly compositional tool-use scenarios with difficulty scaled along both quantitative and structural dimensions, designed to elicit skill abstraction and cross-task reuse. We further propose a lightweight evaluation protocol that enables agents to auto-compose atomic tools into executable Skills, cache and reuse them inside and across tasks, thereby improving efficiency while accumulating a persistent library of reusable skills. Evaluating state-of-the-art agents on SkillCraft, we observe substantial efficiency gains, with token usage reduced by up to 80% by skill saving and reuse. Moreover, success rate strongly correlates with tool composition ability at test time, underscoring compositional skill acquisition as a core capability.

智能体技能复用工具组合效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。