让机器人像人一样思考空间,减少导航时的幻觉错误。
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
- 用双地图双定向模拟人脑空间认知,提升环境理解能力。
- 零样本迁移至真实实验室,性能超越现有最强方法。
- 适合研究具身智能、机器人导航与生物启发式算法的读者。
提升移动机器人的空间感知能力对于实现具身视觉-语言导航(VLN)至关重要。尽管在模拟环境中已取得显著进展,但直接将这些能力迁移到真实场景常导致严重幻觉现象,使机器人丧失有效空间意识。为此,我们提出BrainNav,一种受生物空间认知理论和认知地图理论启发的具身空间认知导航框架。BrainNav融合双地图(坐标图与拓扑图)和双定向(相对方向与绝对方向)策略,支持动态场景捕捉与实时路径规划。其五个核心模块——海马体记忆中枢、视觉皮层感知引擎、顶叶空间构造器、前额叶决策中心和小脑运动执行单元——模拟生物认知功能,减少空间幻觉并增强适应性。在使用Limo Pro机器人的真实世界零样本实验室环境中验证,BrainNav兼容GPT-4,无需微调即优于现有视觉-语言导航连续环境(VLN-CE)的最先进方法。
原文摘要 · Abstract (English)
Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly transferring these capabilities to real-world scenarios often results in severe hallucination phenomena, causing robots to lose effective spatial awareness. To address this issue, we propose BrainNav, a bio-inspired spatial cognitive navigation framework inspired by biological spatial cognition theories and cognitive map theory. BrainNav integrates dual-map (coordinate map and topological map) and dual-orientation (relative orientation and absolute orientation) strategies, enabling real-time navigation through dynamic scene capture and path planning. Its five core modules-Hippocampal Memory Hub, Visual Cortex Perception Engine, Parietal Spatial Constructor, Prefrontal Decision Center, and Cerebellar Motion Execution Unit-mimic biological cognitive functions to reduce spatial hallucinations and enhance adaptability. Validated in a zero-shot real-world lab environment using the Limo Pro robot, BrainNav, compatible with GPT-4, outperforms existing State-of-the-Art (SOTA) Vision-and-Language Navigation in Continuous Environments (VLN-CE) methods without fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。