用在线探索生成高质量动作数据,提升大模型智能体训练效果
LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
- 构建可自主探索任务的交互环境,支持工具调用与实时反馈
- 自动生成数据使模型在基准测试中最高提升49.3%性能
- 几乎无需人工干预,适合快速迭代智能体系统开发
大型行动模型(LAM)在人工智能代理领域潜力巨大,但其训练依赖高质量数据,尤其在涉及规划、调用工具和响应反馈的多步任务中面临挑战。为此,我们提出 LAM SIMULATOR,一个面向智能体任务在线探索与高质量反馈的数据生成框架。该框架包含动态任务生成器、丰富的工具集合以及支持大语言模型(LLM)代理调用工具并接收实时反馈的交互环境,使其能自主探索并解决任务,发现多种解法。由此产生的动作轨迹数据被用于构建高质量训练数据集。在 ToolBench 与 CRMArena 等主流智能体基准上的实验表明,使用本框架自生成数据训练的模型,性能相较原基线最高提升 49.3%。该方法在数据生成阶段几乎无需人工参与,显著提升了 AI 代理研发效率。
原文摘要 · Abstract (English)
Large Action Models (LAMs) for AI Agents offer incredible potential but face challenges due to the need for high-quality training data, especially for multi-steps tasks that involve planning, executing tool calls, and responding to feedback. To address these issues, we present LAM SIMULATOR, a comprehensive framework designed for online exploration of agentic tasks with high-quality feedback. Our framework features a dynamic task query generator, an extensive collection of tools, and an interactive environment where Large Language Model (LLM) Agents can call tools and receive real-time feedback. This setup enables LLM Agents to explore and solve tasks autonomously, facilitating the discovery of multiple approaches to tackle any given task. The resulting action trajectory data are then used to create high-quality training datasets for LAMs. Our experiments on popular agentic benchmarks, ToolBench and CRMArena, highlight the effectiveness of LAM SIMULATOR: models trained with self-generated datasets using our framework achieve significant performance gains, up to a 49.3\% improvement over their original baselines. LAM SIMULATOR requires minimal human input during dataset creation, highlighting LAM SIMULATOR's efficiency and effectiveness in speeding up development of AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。