arXiv:2604.17739cs.LGcs.CL2026-04被引 1

用80亿参数开源模型模拟动态环境,低成本训练工具调用智能体。

Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model

论文配图:Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
图 1 · 摘自论文原文
  • 用80亿参数开源语言模型全程模拟任务生成、用户行为和工具执行
  • 在多数场景超越需外部资源的基线方法,证明小模型也能构建强基准
  • 适合资源有限的研究者,推动工具学习的普惠化与低门槛研究

强化学习已成为训练工具调用智能体的主流范式,通常依赖在线交互环境。现有方法或依赖带真实标注的训练数据,或需高级专有语言模型合成固定环境。本文提出TRUSTEE,一种低成本方法:使用最小可达80亿参数的免费开源语言模型,全程模拟任务生成、用户行为、工具执行及轨迹评估,并结合自适应课程学习机制动态调控任务难度。实验表明,TRUSTEE在多数情况下优于需额外外部资源的基线方法。结果证实,只要设计足够精巧,仅以本地80亿参数语言模型为骨干的仿真环境即可成为工具学习的强大基准。我们希望该范式能推动工具学习的民主化,并激励未来在资源受限条件下开展环境规模化研究。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become a prevalent paradigm for training tool calling agents, which typically requires online interactive environments. Existing approaches either rely on training data with ground truth annotations or require advanced proprietary language models (LMs) to synthesize environments that keep fixed once created. In this work, we propose TRUSTEE, a cost-friendly method for training tool calling agents with dynamic environments fully simulated by free open-source LMs that can be as small as 8B, including task generation, user simulation, tool simulation and trajectory evaluation, paired with an adaptive curriculum learning mechanism that controls task difficulty during training. Our empirical results show that TRUSTEE outperforms baselines which require extra external resources in most cases. These confirm that, with a sufficiently sophisticated design, even simulated environments with a local 8B LM as the backbone could set a strong baseline for tool learning. We hope our proposed paradigm could democratize tool learning and inspire future research on environment scaling with limited resources.

工具学习仿真环境开源模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。