整合感知、规划与预测,提升机器人导航的准确性与长期适应性。
RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction
- 通过多任务协同统一建模感知、规划与预测能力
- 在CHORES-S基准上实现81.1%的成功率,刷新纪录
- 自适应3D历史采样,避免冗余记忆干扰
在语言引导的视觉导航中,智能体需在未见过的环境中根据自然语言指令定位目标物体。为在陌生场景中可靠导航,智能体应具备强大的感知、规划与预测能力。此外,在长期导航过程中重访已探索区域时,智能体可能保留无关且冗余的历史感知信息,导致性能下降。本文提出RoboTron-Nav,一个通过导航与具身问答任务的多任务协作,统一集成感知、规划与预测能力的框架,从而提升导航表现。同时,RoboTron-Nav采用自适应3D感知历史采样策略,高效利用历史观测。借助大语言模型,其可理解多样化指令与复杂视觉场景,生成恰当导航动作。在CHORES-S基准上,RoboTron-Nav达到81.1%的对象目标导航成功率,创下新纪录。
原文摘要 · Abstract (English)
In language-guided visual navigation, agents locate target objects in unseen environments using natural language instructions. For reliable navigation in unfamiliar scenes, agents should possess strong perception, planning, and prediction capabilities. Additionally, when agents revisit previously explored areas during long-term navigation, they may retain irrelevant and redundant historical perceptions, leading to suboptimal results. In this work, we propose RoboTron-Nav, a unified framework that integrates perception, planning, and prediction capabilities through multitask collaborations on navigation and embodied question answering tasks, thereby enhancing navigation performances. Furthermore, RoboTron-Nav employs an adaptive 3D-aware history sampling strategy to effectively and efficiently utilize historical observations. By leveraging large language model, RoboTron-Nav comprehends diverse commands and complex visual scenes, resulting in appropriate navigation actions. RoboTron-Nav achieves an 81.1% success rate in object goal navigation on the $\mathrm{CHORES}$-$\mathbb{S}$ benchmark, setting a new state-of-the-art performance. Project page: https://yvfengzhong.github.io/RoboTron-Nav
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。