让大模型在开放世界中高效使用新工具,提升执行成功率10.8%。
ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution

- 通过主动检索与推理循环实现动态工具调用
- 端到端执行成功率提升10.8%,优于现有方法
- 适合需要持续学习新工具的智能代理场景
大型语言模型(LLMs)可通过调用外部工具增强问题求解能力。然而,在工具库海量且不断演化的开放世界场景中,依赖静态嵌入检索或工具参数记忆的方法,难以对齐用户意图与工具语义,或泛化至未见过的工具,导致工具检索与执行准确率不足。为此,我们提出ToolOmni,一种统一的智能体框架,通过推理循环中的主动检索与具身执行,使LLM具备开放世界工具使用能力。首先,构建冷启动多轮交互数据集,通过监督微调(SFT)赋予基础智能体能力;其次,提出基于解耦多目标GRPO算法的开放世界工具学习方法,在线优化模型的工具检索准确率与执行有效性。大量实验表明,ToolOmni在检索与执行上均达到领先性能,端到端执行成功率相较强基线提升+10.8%,展现出卓越的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) enhance their problem-solving capability by utilizing external tools. However, in open-world scenarios with massive and evolving tool repositories, existing methods relying on static embedding retrieval or parameter memorization of tools struggle to align user intent with tool semantics or generalize to unseen tools, respectively, leading to suboptimal accuracy of open-world tool retrieval and execution. To address these, we present ToolOmni, a unified agentic framework that enables LLMs for open-world tool use by proactive retrieval and grounded execution within a reasoning loop. First, we construct a cold-start multi-turn interaction dataset to instill foundational agentic capabilities via Supervised Fine-Tuning (SFT). Then, we introduce open-world tool learning based on a Decoupled Multi-Objective GRPO algorithm, which simultaneously optimizes LLMs for both tool retrieval accuracy and execution efficacy in online environments. Extensive experiments demonstrate that ToolOmni achieves state-of-the-art performance both in retrieval and execution, surpassing strong baselines by a significant margin of +10.8% in end-to-end execution success rate, while exhibiting exceptional robustness and generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。