首次从神经辐射场中直接提取物体间语义关系,实现开放词汇场景理解。
RelationField: Relate Anything in Radiance Fields
- 用射线对表示物体间关系,扩展神经辐射场隐式查询能力
- 通过多模态大模型蒸馏知识,支持复杂开放词汇关系建模
- 在3D场景图生成和关系引导分割任务上达最优性能
神经辐射场是新兴的三维场景表征方法,近年已通过蒸馏视觉-语言模型的开放词汇特征扩展至场景理解。但现有方法主要聚焦于以对象为中心的表示,仅支持对象分割或检测,而对物体间语义关系的理解仍基本空白。为此,我们提出RelationField,首个直接从神经辐射场中提取物体间关系的方法。RelationField将物体间关系表示为神经辐射场内的射线对,有效扩展了其公式以支持隐式关系查询。为让RelationField学习复杂、开放词汇的关系,关系知识由多模态大模型蒸馏而来。为评估该方法,我们解决了开放词汇3D场景图生成与关系引导实例分割任务,在两项任务上均取得当前最佳性能。
原文摘要 · Abstract (English)
Neural radiance fields are an emerging 3D scene representation and recently even been extended to learn features for scene understanding by distilling open-vocabulary features from vision-language models. However, current method primarily focus on object-centric representations, supporting object segmentation or detection, while understanding semantic relationships between objects remains largely unexplored. To address this gap, we propose RelationField, the first method to extract inter-object relationships directly from neural radiance fields. RelationField represents relationships between objects as pairs of rays within a neural radiance field, effectively extending its formulation to include implicit relationship queries. To teach RelationField complex, open-vocabulary relationships, relationship knowledge is distilled from multi-modal LLMs. To evaluate RelationField, we solve open-vocabulary 3D scene graph generation tasks and relationship-guided instance segmentation, achieving state-of-the-art performance in both tasks. See the project website at https://relationfield.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。