arXiv:2601.10168cs.CVcs.AI2026-01中稿 · ECCV被引 4

用重拍视角提升3D场景图语义一致性,让机器人看得更准。

RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation

  • 通过重拍视角对比,量化物体语义不确定性。
  • 低不确定物体作锚点,检索可靠知识修复高不确定预测。
  • 在三个基准和真实机器人测试中显著提升准确率与召回率。

开放词汇的3D场景图(3DSG)可通过结构化语义表示增强机器人下游任务,但现有构建方法在遮挡和视角受限条件下,因跨图像聚合噪声导致语义不一致。为此,我们提出RAG-3DSG,引入重拍引导的不确定性估计:通过比较原始有限视角与重拍最优视角间的语义一致性,量化每个图中物体的潜在语义模糊性。基于此量化结果,设计对象级检索增强生成(RAG),以低不确定性物体为语义锚点,检索更可靠的上下文知识,驱动视觉语言模型修正不确定物体的预测,优化最终3DSG。在三个挑战性基准和真实机器人实验中,RAG-3DSG均实现更高召回率与精确率,有效缓解语义噪声,为机器人任务提供高度可靠的场景表征。

原文摘要 · Abstract (English)

Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image aggregation under occlusions and constrained viewpoints. To mitigate the impact of such inconsistency, we propose RAG-3DSG, which introduces re-shot guided uncertainty estimation. By measuring the semantic consistency between original limited viewpoints and re-shot optimal viewpoints, this method quantifies the underlying semantic ambiguity of each graph object. Based on this quantification, we devise an Object-level Retrieval-Augmented Generation (RAG) that leverages low-uncertainty objects as semantic anchors to retrieve more reliable contextual knowledge, enabling a Vision-Language Model to rectify the predictions of uncertain objects and optimize the final 3DSG. Extensive evaluations across three challenging benchmarks and real-world robot trials demonstrate that RAG-3DSG achieves superior recall and precision, effectively mitigating semantic noise to provide highly reliable scene representations for robotics tasks.

3D场景图视觉语言模型机器人感知检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。