arXiv:2605.10118cs.RO2026-05中稿 · ICML

用抽象物理环境训练机器人导航,提升真实世界适应力。

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation

论文配图:Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation
图 1 · 摘自论文原文
  • 在简化物理环境中预演导航计划,替代真实感模拟
  • 在A-EQA数据集上达成53.21%的成功率,比基线高9.7%
  • 适合需跨域迁移的机器人导航研究者

视觉语言模型虽具强大推理能力,但在具身导航中受限于开放世界视觉与机器人控制数据的缺乏。尽管模拟器可低成本收集数据,但依赖真实感仿真常导致策略难以迁移。为此,我们提出SAGE框架,让智能体在物理约束的语义抽象环境中学习,类比人类心理模拟:先在简化环境中演练计划,再执行。SAGE包含三个协同阶段:(1) 创生阶段,构建多样且受物理约束的语义环境以启动经验;(2) 演化阶段,通过强化学习蒸馏经验,采用新颖的非对称自适应裁剪机制稳定更新;(3) 导航阶段,将抽象策略映射至真实世界控制。实验表明,SAGE显著提升规划辅助的具身导航性能,在A-EQA数据集上达到53.21%的LLM-Match成功率(较基线提升9.7%),并在物理室内机器人部署中展现良好迁移能力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated exceptional general reasoning capabilities. However, their performance in embodied navigation remains hindered by a scarcity of aligned open-world vision and robot control data. Despite simulators providing a cost-effective alternative for data collection, the inherent reliance on photorealistic simulations often limits the transferability of learned policies. To this end, we propose \textit{\textbf{S}andbox-\textbf{A}bstracted \textbf{G}rounded \textbf{E}xperience} (\textbf{\textit{SAGE}}), a framework that enables agents to learn within a physics-grounded semantic abstraction rather than a photorealistic simulation, mimicking the human capacity for mental simulation where plans are rehearsed in simplified physics abstractions before execution. \textit{SAGE} system operates via three synergistic phases: (1) \textit{Genesis}: constructing diverse, physics-constrained semantic environments to bootstrap experience; (2) \textit{Evolution}: distilling experiences through Reinforcement Learning (RL), utilizing a novel asymmetric adaptive clipping mechanism to stabilize updates; (3) \textit{Navigation}: bridging the abstract policy to open-world control. We demonstrate that \textit{SAGE} significantly improves planner-assisted embodied navigation, achieving a 53.21\% LLM-Match Success Rate on A-EQA (+9.7\% over baseline), while showing encouraging transfer to physical indoor robot deployment.

具身导航物理抽象强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。