arXiv:2605.20823cs.CV2026-05

用视觉几何线索补全3D场景关系,让模型更准地识别未标注的关系。

RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses

论文配图:RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
图 1 · 摘自论文原文
  • 利用视觉几何线索构建关系证据,判断物体间关系是否可观察。
  • 在3DSSG/3RScan数据集上,未见关系识别率提升,幻觉和冗余描述减少。
  • 适合需要高可靠性3D场景理解的开放词汇应用,如机器人导航与交互。

开放词汇3D场景图生成旨在用灵活的自然语言谓词描述物体实例及其关系。核心难点不仅在于词汇扩展,更在于监督可靠性:现有3D场景图数据集中的关系标注具有选择性,许多有效物体对关系未被标注。我们提出RelWitness框架,从带姿态的RGB-D序列中进行开放词汇3D场景图生成,面对不完整的关系监督。关键概念是关系见证(relation witness):一种使关系在场景中可观察的视觉-几何线索。支撑关系需接触与垂直排序;包含关系需封闭;邻近关系需度量接近;朝向关系需正对方向;稳定关系应在双物体可见时持续存在。RelWitness从RGB视图、深度图、重建3D几何、角色敏感文本、物体先验空视图及多视角一致性构建关系见证记录。视觉-几何见证验证器将未标注关系候选分类为已验证的缺失正例、可靠负例或不确定未标记情况。基于见证的正-未标记目标则在不将每个缺失标签视为负例的前提下学习。我们还引入见证一致解码和RGB-D缺失关系审计协议。在3DSSG/3RScan及扫描网衍生的开放词汇划分上的模拟论文规划实验显示预期行为:未见关系识别能力提升,见证精度提高,幻觉减少,冗余关系短语降低。所有数值结果为规划值,提交前需替换为复现测量结果。

原文摘要 · Abstract (English)

Open-vocabulary 3D scene graph generation seeks to describe object instances and their relations with flexible natural-language predicates. The central difficulty is not only vocabulary expansion, but supervision reliability: relation annotations in 3D scene graph datasets are selective, and many valid object-pair relations are unannotated. We propose RelWitness, a framework for open-vocabulary 3D scene graph generation from posed RGB-D sequences under incomplete relation supervision. The key concept is a relation witness: a concrete visual-geometric cue that makes a relation observable in the captured scene. Support relations require contact and vertical ordering; containment requires enclosure; proximity requires metric closeness; orientation requires facing direction; and stable relations should persist across views where both objects are visible. RelWitness constructs relation witness records from RGB views, depth maps, reconstructed 3D geometry, role-sensitive text, object-prior null views, and multi-view consistency. A visual-geometric witness verifier assigns unannotated relation candidates to verified missing positives, reliable negatives, or uncertain unlabeled cases. A witness-guided positive-unlabeled objective then learns from incomplete annotations without turning every missing label into a negative. We further introduce witness-consistent decoding and an RGB-D missing-relation audit protocol. Simulated manuscript-planning experiments on 3DSSG/3RScan and ScanNet-derived open-vocabulary splits show the intended behavior: improved unseen-relation recognition, higher witness precision, lower hallucination, and reduced redundant relation phrases. All numerical results are planning values and must be replaced by reproduced measurements before submission

3D场景图开放词汇关系检测几何线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。