用程序化生成交互式工具数据,提升智能体工具使用能力
Procedural Environment Generation for Tool-Use Agents
- 设计随机世界管道,自动生成可交互且可组合的工具数据
- 在NESTFUL数据集上两项指标刷新最佳纪录
- 合成数据量越多,下游性能越强,适合强化学习研究者
尽管大语言模型工具使用智能体的潜力已引发大量研究,但工具使用训练数据的构建仍是开放问题,尤其适用于在线强化学习训练。现有合成数据生成方法往往缺乏交互性或组合性。我们提出RandomWorld,一个用于程序化生成交互式工具和组合性工具使用数据的流水线。实验表明,基于SFT和RL在合成RandomWorld数据上训练的模型,在多个工具使用基准上表现更优,并在NESTFUL数据集的两项指标上达到新最优。进一步实验显示,下游性能随RandomWorld生成训练数据量增加而提升,为完全使用合成数据实现性能突破提供了可能。
原文摘要 · Abstract (English)
Although the power of LLM tool-use agents has ignited a flurry of recent research in this area, the curation of tool-use training data remains an open problem$-$especially for online RL training. Existing approaches to synthetic tool-use data generation tend to be non-interactive, and/or non-compositional. We introduce RandomWorld, a pipeline for the procedural generation of interactive tools and compositional tool-use data. We show that models tuned via SFT and RL on synthetic RandomWorld data improve on a range of tool-use benchmarks, and set the new SoTA for two metrics on the NESTFUL dataset. Further experiments show that downstream performance scales with the amount of RandomWorld-generated training data, opening up the possibility of further improvement through the use of entirely synthetic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。