用世界模型降低高成本任务中的试错成本,提升智能体表现
World Models as an Intermediary between Agents and the Real World
- 以世界模型作为智能体与真实世界的中介,模拟环境动态与奖励
- 显著改善长周期任务的样本效率和离策略学习问题
- 适用于机器人、科学实验等高成本领域,适合研究复杂决策系统者
使用强化学习训练的大语言模型智能体在游戏、数学和编程等低成本环境中已达到超人水平。然而,这些成功尚未扩展到高成本复杂领域,如机器人物理操作、机器学习工程耗时以及科学实验资源消耗。实现更高性能的关键瓶颈在于执行动作获取奖励信号的高昂代价。本文提出应将世界模型作为智能体与真实世界之间的中介。我们探讨了世界模型作为动态、奖励与任务分布的建模工具,如何克服高成本行动带来的根本性障碍,包括极端离策略学习和长周期任务中的样本效率低下。此外,我们展示了世界模型如何在机器学习工程、计算机使用、机器人及人工智能科学等多个领域为智能体提供丰富而关键的学习信号。最后,我们指出了构建世界模型的挑战,并针对数据集构建、架构设计、规模扩展和评估方法提出了可操作建议。
原文摘要 · Abstract (English)
Large language model (LLM) agents trained using reinforcement learning has achieved superhuman performance in low-cost environments like games, mathematics, and coding. However, these successes have not translated to complex domains where the cost of interaction is high, such as the physical cost of running robots, the time cost of ML engineering, and the resource cost of scientific experiments. The true bottleneck for achieving the next level of agent performance for these complex and high-cost domains lies in the expense of executing actions to acquire reward signals. To address this gap, this paper argues that we should use world models as an intermediary between agents and the real world. We discuss how world models, viewed as models of dynamics, rewards, and task distributions, can overcome fundamental barriers of high-cost actions such as extreme off-policy learning and sample inefficiency in long-horizon tasks. Moreover, we demonstrate how world models can provide critical and rich learning signals to agents across a broad set of domains, including machine learning engineering, computer use, robotics, and AI for science. Lastly, we identify the challenges of building these world models and propose actionable items along dataset curation, architecture design, scaling, and evaluation of world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。