提升点云实例分割中关系建模能力,增强特征表示与查询间关联
Relation3D: Enhancing Relation Modeling for Point Cloud Instance Segmentation

- 设计自适应超点聚合与对比学习优化模块,强化场景特征表达
- 引入关系感知自注意力机制,融合位置与几何关系提升查询建模
- 在多个主流数据集上超越现有方法,适合3D实例分割研究者参考
3D实例分割旨在预测场景中一组物体实例,以二值前景掩码和对应语义标签表示。当前基于Transformer的方法因结构简洁且性能优越而受到关注,但主要依赖掩码注意力建模场景特征与查询特征间的外部关系,缺乏对场景特征内部关系及查询特征间关系的有效建模。针对此问题,我们提出Relation3D:通过自适应超点聚合模块与对比学习引导的超点精化模块,更好地表征超点特征(场景特征),并利用对比学习指导其更新;此外,关系感知自注意力机制将位置与几何关系融入自注意力过程,增强查询间关系建模能力。在ScanNetV2、ScanNet++、ScanNet200和S3DIS等多个数据集上的大量实验表明,Relation3D表现优异。
原文摘要 · Abstract (English)
3D instance segmentation aims to predict a set of object instances in a scene, representing them as binary foreground masks with corresponding semantic labels. Currently, transformer-based methods are gaining increasing attention due to their elegant pipelines and superior predictions. However, these methods primarily focus on modeling the external relationships between scene features and query features through mask attention. They lack effective modeling of the internal relationships among scene features as well as between query features. In light of these disadvantages, we propose \textbf{Relation3D: Enhancing Relation Modeling for Point Cloud Instance Segmentation}. Specifically, we introduce an adaptive superpoint aggregation module and a contrastive learning-guided superpoint refinement module to better represent superpoint features (scene features) and leverage contrastive learning to guide the updates of these features. Furthermore, our relation-aware self-attention mechanism enhances the capabilities of modeling relationships between queries by incorporating positional and geometric relationships into the self-attention mechanism. Extensive experiments on the ScanNetV2, ScanNet++, ScanNet200 and S3DIS datasets demonstrate the superior performance of Relation3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。