多智能体协作构建开放词汇的动态城市场景图谱
Collaborative Dynamic 3D Scene Graphs for Open-Vocabulary Urban Scene Understanding
- 多智能体融合视觉与激光雷达数据,协同构建场景图谱
- 在真实数据上实现比单智能体更高精度的地图与物体预测
- 支持开放词汇语义层次结构,适合真实城市环境研究
地图构建与场景表征是移动机器人可靠规划与导航的基础。尽管基于体素网格的纯几何地图可实现通用导航,但在动态大尺度环境中获取实时、语义丰富的表征仍具挑战。本文提出CURB-OSG,一种开放词汇的动态3D场景图引擎,通过多智能体协作对城市驾驶场景进行分层分解。该方法融合多个感知智能体(初始位姿未知)的相机与激光雷达观测,生成比单智能体更精确的地图,并构建统一的开放词汇语义层级。不同于依赖真值位姿或仅在仿真中评估的方法,CURB-OSG降低了实际应用约束。我们在牛津雷达机器人汽车数据集的多时段真实多智能体传感器数据上验证了其性能,展示了多智能体协作带来的地图与物体预测精度提升,以及环境分割能力。代码与补充材料已公开。
原文摘要 · Abstract (English)
Mapping and scene representation are fundamental to reliable planning and navigation in mobile robots. While purely geometric maps using voxel grids allow for general navigation, obtaining up-to-date spatial and semantically rich representations that scale to dynamic large-scale environments remains challenging. In this work, we present CURB-OSG, an open-vocabulary dynamic 3D scene graph engine that generates hierarchical decompositions of urban driving scenes via multi-agent collaboration. By fusing the camera and LiDAR observations from multiple perceiving agents with unknown initial poses, our approach generates more accurate maps compared to a single agent while constructing a unified open-vocabulary semantic hierarchy of the scene. Unlike previous methods that rely on ground truth agent poses or are evaluated purely in simulation, CURB-OSG alleviates these constraints. We evaluate the capabilities of CURB-OSG on real-world multi-agent sensor data obtained from multiple sessions of the Oxford Radar RobotCar dataset. We demonstrate improved mapping and object prediction accuracy through multi-agent collaboration as well as evaluate the environment partitioning capabilities of the proposed approach. To foster further research, we release our code and supplementary material at https://ov-curb.cs.uni-freiburg.de.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。