arXiv:2409.18108cs.RO2024-09被引 30

用移动机器人逐步构建带语义的3D地图,支持任意物体查询。

Language-Embedded Gaussian Splats (LEGS): Incrementally Building Room-Scale Representations with a Mobile Robot

  • 基于多相机与增量式捆绑调整,边走边学构建3D场景
  • 对开放词汇和长尾物体查询准确率达66%
  • 训练速度比同类系统快3.5倍,适合实际部署

构建语义3D地图对在办公室、仓库、商店和家庭中查找目标物体具有重要价值。本文提出一种增量式映射系统,构建名为语言嵌入高斯点(LEGS)的统一3D场景表示,同时编码外观与语义信息。LEGS在机器人巡游环境中在线训练,支持开放词汇物体查询定位。我们在4个室级场景上评估,通过查询场景中的物体来检验其语义捕捉能力。与LERF相比,两者物体查询成功率相近,但LEGS训练速度快3.5倍以上。结果表明,多相机设置与增量式捆绑调整可提升受限轨迹下的视觉重建质量,且LEGS能以最高66%的准确率定位开放词汇及长尾物体查询。

原文摘要 · Abstract (English)

Building semantic 3D maps is valuable for searching for objects of interest in offices, warehouses, stores, and homes. We present a mapping system that incrementally builds a Language-Embedded Gaussian Splat (LEGS): a detailed 3D scene representation that encodes both appearance and semantics in a unified representation. LEGS is trained online as a robot traverses its environment to enable localization of open-vocabulary object queries. We evaluate LEGS on 4 room-scale scenes where we query for objects in the scene to assess how LEGS can capture semantic meaning. We compare LEGS to LERF and find that while both systems have comparable object query success rates, LEGS trains over 3.5x faster than LERF. Results suggest that a multi-camera setup and incremental bundle adjustment can boost visual reconstruction quality in constrained robot trajectories, and suggest LEGS can localize open-vocabulary and long-tail object queries with up to 66% accuracy.

3D重建语义地图机器人感知开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。