解决3D实例分割中大物体过分割问题,提升掩码精度
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
- 多尺度特征+双注意力机制,捕捉更丰富的空间语义
- 引入框查询与框正则化,缓解掩码预测不可靠问题
- 在ScanNetV2等数据集上超越现有方法,适合高精度场景理解
近期基于Transformer并结合超点的3D实例分割方法广泛应用,但常面临大物体过分割问题,且由超点掩码预测不准确进一步加剧。为此,我们提出MSTA3D框架,利用多尺度特征表示并引入双注意力机制以有效捕获特征。此外,MSTA3D融合框查询与框正则化,在语义查询外提供互补的空间约束。在ScanNetV2、ScanNet200和S3DIS数据集上的实验表明,该方法优于当前最先进的3D实例分割方法。
原文摘要 · Abstract (English)
Recently, transformer-based techniques incorporating superpoints have become prevalent in 3D instance segmentation. However, they often encounter an over-segmentation problem, especially noticeable with large objects. Additionally, unreliable mask predictions stemming from superpoint mask prediction further compound this issue. To address these challenges, we propose a novel framework called MSTA3D. It leverages multi-scale feature representation and introduces a twin-attention mechanism to effectively capture them. Furthermore, MSTA3D integrates a box query with a box regularizer, offering a complementary spatial constraint alongside semantic queries. Experimental evaluations on ScanNetV2, ScanNet200 and S3DIS datasets demonstrate that our approach surpasses state-of-the-art 3D instance segmentation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。