从二元关系扩展到多元关系,提升3D物体定位的准确性
B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding
- 构建从二元到多元关系的渐进式学习框架
- 在ReferIt3D和ScanRefer上显著优于现有方法
- 适合需要精细场景理解的机器人应用
使用自然语言定位3D物体对机器人场景理解至关重要。描述常涉及多个空间关系以区分相似物体,导致多模态对齐困难。现有方法仅建模成对物体间的关系,忽略了多模态关系理解中多元组合的全局感知意义。为此,我们提出一种新颖的渐进式关系学习框架用于3D物体定位,将关系学习从二元扩展至多元,以识别全局匹配指代描述的视觉关系。由于训练数据中缺乏目标物体的具体标注,我们设计了分组监督损失以促进多元关系学习。在包含多元关系的场景图中,采用融合注意力机制的多模态网络进一步精确定位目标物体。在ReferIt3D和ScanRefer基准上的实验与消融研究证明,该方法优于当前最先进水平,验证了多元关系感知在3D定位中的优势。
原文摘要 · Abstract (English)
Localizing 3D objects using natural language is essential for robotic scene understanding. The descriptions often involve multiple spatial relationships to distinguish similar objects, making 3D-language alignment difficult. Current methods only model relationships for pairwise objects, ignoring the global perceptual significance of n-ary combinations in multi-modal relational understanding. To address this, we propose a novel progressive relational learning framework for 3D object grounding. We extend relational learning from binary to n-ary to identify visual relations that match the referential description globally. Given the absence of specific annotations for referred objects in the training data, we design a grouped supervision loss to facilitate n-ary relational learning. In the scene graph created with n-ary relationships, we use a multi-modal network with hybrid attention mechanisms to further localize the target within the n-ary combinations. Experiments and ablation studies on the ReferIt3D and ScanRefer benchmarks demonstrate that our method outperforms the state-of-the-art, and proves the advantages of the n-ary relational perception in 3D localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。