arXiv:2503.24388cs.AIcs.CL2025-03被引 5

让智能体先推理再想象,提升决策效率与泛化能力。

RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy

  • 将推理与想象统一在端到端策略中,通过联合学习建模动作与环境动态关系。
  • 样本效率提升17倍以上,且在多个任务上展现更强泛化能力。
  • 适合需要高鲁棒性与自纠错能力的通用智能体研究者使用。

在复杂开放世界环境中,行动前的推理和对潜在结果的想象(即世界模型)对具身智能体至关重要。然而,以往工作要么仅包含其中一种能力,要么将多个专用模型集成到系统中,限制了策略的学习效率与泛化性能。为此,本文首次提出在端到端通用策略中协同融合推理与想象,命名为RIG。为实现端到端训练,构建了一条数据流水线,逐步整合并丰富从现有智能体轨迹中收集的想象与推理内容。推理与下一图像生成的联合学习显式建模了推理、动作与环境动态之间的内在关联,相比先前方法,样本效率提升超过17倍,并展现出更优的泛化能力。推理时,RIG先推断下一步动作,生成潜在动作,再预测动作结果,使智能体能在真实执行前基于想象进行复核与自我修正。实验表明,推理与想象的协同不仅增强了通用策略的鲁棒性、泛化性与互操作性,还支持测试时扩展以进一步提升性能。

原文摘要 · Abstract (English)

Reasoning before action and imagining potential outcomes (i.e., world models) are essential for embodied agents operating in complex open-world environments. Yet, prior work either incorporates only one of these abilities in an end-to-end agent or integrates multiple specialized models into an agent system, limiting the learning efficiency and generalization of the policy. Thus, this paper makes the first attempt to synergize Reasoning and Imagination in an end-to-end Generalist policy, termed RIG. To train RIG in an end-to-end manner, we construct a data pipeline that progressively integrates and enriches the content of imagination and reasoning in the trajectories collected from existing agents. The joint learning of reasoning and next image generation explicitly models the inherent correlation between reasoning, action, and dynamics of environments, and thus exhibits more than $17\times$ sample efficiency improvements and generalization in comparison with previous works. During inference, RIG first reasons about the next action, produces potential action, and then predicts the action outcomes, which offers the agent a chance to review and self-correct based on the imagination before taking real actions. Experimental results show that the synergy of reasoning and imagination not only improves the robustness, generalization, and interoperability of generalist policy but also enables test-time scaling to enhance overall performance.

通用策略推理想象端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。