多模态感知让机器人长距离导航更可靠,能识路、避障、懂社交规则。
MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding
- 融合激光、图像与视觉语言模型,实现环境理解
- 在真实道路场景中可通行性提升10%
- 适合需要复杂室外导航的机器人系统
我们提出MOSU,一种新型自主长距离机器人导航系统,通过多模态感知与道路场景理解提升移动机器人的全局导航能力。该系统整合几何、语义与上下文信息,实现全面的场景认知。高层路径规划采用GPS与QGIS地图路由,局部导航则通过多模态轨迹生成进行优化。轨迹生成利用多模态数据:基于LiDAR的几何信息用于精确避障,图像语义分割评估可通行性,视觉语言模型(VLMs)捕捉社会上下文,使机器人在复杂环境中遵守社交规范。这种多模态融合提升了场景理解与可通行性,使机器人适应多样化的户外条件。我们在真实道路环境进行测试,并在GND数据集上评估,结果显示在可通行地形上可通行性提升10%,同时保持与现有全局导航方法相当的导航距离。
原文摘要 · Abstract (English)
We present MOSU, a novel autonomous long-range navigation system that enhances global navigation for mobile robots through multimodal perception and on-road scene understanding. MOSU addresses the outdoor robot navigation challenge by integrating geometric, semantic, and contextual information to ensure comprehensive scene understanding. The system combines GPS and QGIS map-based routing for high-level global path planning and multi-modal trajectory generation for local navigation refinement. For trajectory generation, MOSU leverages multi-modalities: LiDAR-based geometric data for precise obstacle avoidance, image-based semantic segmentation for traversability assessment, and Vision-Language Models (VLMs) to capture social context and enable the robot to adhere to social norms in complex environments. This multi-modal integration improves scene understanding and enhances traversability, allowing the robot to adapt to diverse outdoor conditions. We evaluate our system in real-world on-road environments and benchmark it on the GND dataset, achieving a 10% improvement in traversability on navigable terrains while maintaining a comparable navigation distance to existing global navigation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。