构建智能体探索能力的新基准,用积木搭建测试真实交互学习。
BuilderBench: The Building Blocks of Intelligent Agents
- 设计物理仿真环境与50+结构任务,驱动智能体通过试错学习。
- 主流大模型与强化学习算法均无法完成复杂搭建任务。
- 适合研究具身智能、开放探索与长程规划的学者使用。
当前AI模型主要依赖模仿和精炼,难以突破已有数据的局限。为解决新问题,智能体需通过交互探索积累技能。如何构建可扩展的学习机制仍是重大挑战。本文提出BuilderBench,一个聚焦开放式探索的智能体训练基准。该基准要求智能体在物理仿真环境中使用积木搭建各种结构。系统包含(1)机器人与多种物理积木的交互模拟器,(2)超过50个精心设计的目标结构任务,涵盖物理理解、数学推理与长程规划能力。智能体在每轮开始时获知目标结构,可通过多轮交互实验并学习搭建技能。完成任务依赖于非语言的具身推理,需尝试不同策略并组合优化。我们在多个前沿基于语言模型的智能体及从零开始的强化学习算法上进行实验,发现它们均无法解决任何非平凡任务。分析揭示了现有模型缺乏有效探索能力。
原文摘要 · Abstract (English)
Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set by existing data. To solve novel problems, agents should acquire skills by exploring and learning through experience. Finding a scalable learning mechanism for developing agents that learn through interaction remains a major open problem. In this work, we introduce BuilderBench, a benchmark to accelerate research into agent training that centers open-ended exploration. BuilderBench requires agents to learn how to build any structure using blocks. BuilderBench is equipped with (1) a simulator of a robot interacting with various physical blocks, and (2) a task-suite with over 50 diverse target structures that are carefully curated to test an understanding of physics, mathematics, and long-horizon planning. Agents are provided with a target structure at the start, and can interact with the environment for multiple episodes to experiment and learn various skills for building the structure. Solving these tasks requires \emph{embodied reasoning} in a way that is not reflected in words but rather in actions, experimenting with different strategies and piecing them together. Our experiments with multiple state-of-the-art frontier language model based agents and tabula rasa reinforcement learning algorithms show that these agents cannot solve any of the non-trivial tasks in the BuilderBench. Our analysis throws light on the lack of exploration abilities in these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。