用语义地图引导机器人自主探索,实现高效精准导航。
RoboAtlas: Contextual Active SLAM

- 通过上下文多臂赌博机动态平衡几何探索与语义推理。
- 在超1800平米场景中完成100%任务成功率,语义实例达3万。
- 小模型也能超越大模型基线,凸显语义地图的关键作用。
我们提出RoboAtlas,一种基于可扩展3D语义地图系统OpenRoboVox的上下文主动式同步定位与建图框架。该框架通过上下文多臂赌博机,融合前沿探索、全局语义地图推理和以我为中心的视觉语言模型(VLM)推理,随场景理解提升逐步从探索转向语义引导导航。我们在仿真环境及单位制Go2机器人的真实大规模环境中进行了评估,覆盖面积超过1800平方米,共映射约3万个语义实例,任务成功率达到100%。在GOAT-Bench 'Val Unseen'基准上,RoboAtlas使用GPT-4o达到90.6%的最高报告成功率,较最强基线提升17.8个百分点。即使使用更小的Qwen2.5-VL-7B模型,仍实现88.8%的成功率,优于所有采用GPT-4o的基线,证明了语义地图所提供的信息价值远超单纯替换基础模型。结果表明,将基础模型与大规模3D语义地图结合,可实现鲁棒高效的上下文主动式SLAM。
原文摘要 · Abstract (English)
We present RoboAtlas, a contextual Active SLAM framework that adaptively balances geometric exploration and semantic reasoning using a scalable 3D semantic mapping system, OpenRoboVox. RoboAtlas integrates frontier exploration, global semantic-map reasoning, and egocentric VLM-based reasoning through a contextual multi-armed bandit that transitions from exploration to semantically guided navigation as scene understanding improves. We evaluate the system in simulation and on a Unitree Go2 robot in large-scale real-world environments exceeding 1800 m2 with approx. 30k mapped semantic instances, achieving a 100% task success rate. On the GOAT-Bench "Val Unseen" benchmark, RoboAtlas achieves state-of-the-art performance with highest reported success rate (SR) of 90.6%, using GPT-4o, improving over the strongest prior baseline by 17.8 percentage points in SR. Using the much smaller Qwen2.5-VL-7B model, it still achieves 88.8% SR, outperforming all baselines using GPT-4o in SR, and revealing the importance of the information gained by our semantic mapping framework over simply replacing the underlying foundation model. The results demonstrate that grounding foundation models with large-scale 3D semantic maps enables robust and efficient contextual Active SLAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。