arXiv:2607.16961cs.AI2026-07

提出工具发现新框架,揭示大模型越做越难发现工具

Lomekwi: Resource-Bounded Tool Discovery in LLM Agents

  • 区分工具使用与发现,拆解为好奇心、识别力和效率三部分
  • 发现识别能力随模型增大反而下降,且在组合游戏中验证
  • 适用于真实任务环境,对智能体工具探索研究有重要启发

现有工具使用基准仅报告复杂多步任务的单一成功率。受认知科学启发,我们区分工具使用与工具发现,并将后者分解为好奇心(发现构建工具所需部件)、识别力(发现创建工具的过程)和效率(创造后使用工具的能力)。该框架可应用于现有发现任务,如Voyager。我们发现识别力随模型规模增大而反向缩放,并引入一类组合游戏加以验证。此外,在模拟真实任务的独立环境中也观察到类似反向缩放现象。

原文摘要 · Abstract (English)

Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distinguish tool use from tool discovery and decompose the latter into curiosity (the model's ability to discover the parts needed to build the tool), recognition (the model's ability to discover the process of creating the tool), and efficiency (the model's use of the tool after creation). We show that this framework can be applied to existing discovery tasks, such as Voyager. In addition, we provide evidence that recognition inversely scales with model size, and we introduce and analyze a class of combinatorial games that demonstrates this. We further observe inverse scaling in a separate environment designed to emulate real-world tasks.

工具发现大模型行为智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。