arXiv:2609.04607cs.RO2026-09

户外机器人用3D场景图可实现近70%的导航寻物成功率,但路径效率不足。

Open-Set 3D Scene Graphs for Field Robotics: An Outdoor Case Study

论文配图:Open-Set 3D Scene Graphs for Field Robotics: An Outdoor Case Study
图 1 · 摘自论文原文
  • 基于视觉语言模型生成点云语义嵌入,构建分层3D场景图。
  • 多数据集测试显示点嵌入中异常点占比超30%,区域理解F1仅0.359。
  • 支持千米级地图压缩至600MB内,适合实际部署,但路径规划仍需优化。

三维场景图(3DSGs)作为支持高级机器人推理的几何精确、语义丰富、层次化的通用地图,近年来备受关注。然而,其在真实户外环境中的表现仍不清晰,尤其与开放集视觉语言模型(VLMs)结合时。本文以Terra 3DSG为案例,分析五个户外机器人数据集中的共性组件,评估语义点嵌入、节点图导航、区域级理解及内存占用。引入新一致性指标,检验重复遍历中图结构与语义的稳定性。结果表明,所有数据集中约30%的点存在异常或多重语义模式,异常比例超过0.1;导航寻物成功率达70%左右,但路径效率平均仅为66%最优路径水平,受可达性失败和路由低效影响;复杂自然环境中区域理解能力弱,平均F1分数仅约0.359。总体而言,户外3DSGs可在多公里轨迹下保持小于600MB的紧凑表示,并具备相对稳定的结构,但仍面临多重语义处理、融合可达性、提升高层区域理解等挑战。

原文摘要 · Abstract (English)

Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounded, semantically informed, hierarchical general-purpose maps to support high-level robotic reasoning. However, the behavior of 3DSGs in real-world outdoor deployments remains poorly understood, particularly when combined with open-set vision-language models (VLMs). In this field report, we analyze the components common to most 3DSG representations across five outdoor robotic datasets to characterize challenges that arise in complex outdoor environments. Using the recently proposed Terra 3DSG as a case study, we investigate semantic point embeddings, place-node graph navigation, region-level understanding, and memory size across the five diverse datasets. We additionally introduce novel consistency metrics to evaluate whether semantic and structural graph properties remain stable across repeated traversals of the same environment. Our analysis reveals that outliers and multiple modes are common in VLM point embeddings across all tested datasets with outlier ratios above $0.1$ for around $30\%$ of points. We demonstrate the feasibility of outdoor 3DSGs for navigation-based object retrieval, achieving success rates near $70\%$, though performance is limited by traversability failures and inefficient routing, with trajectories averaging approximately $66\%$ suboptimal path efficiency. Region-level understanding remains challenging in complex natural environments, with low average F1 scores around $0.359$. Overall, our results show that outdoor 3DSGs can maintain compact (less than $600$MB for multi-kilometer trajectories) and relatively consistent large-scale environment representations, while highlighting open challenges in handling multiple semantic modes, incorporating traversability into graph structures, and improving higher-level region understanding.

3D场景图户外机器人语义导航视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。