用灵活视图集和显式关联建模,提升3D形状理解性能
VSFormer: Mining Correlations in Flexible View Set for Multi-view 3D Shape Understanding
- 将多视图构造成无序集合,打破固定关系假设
- 通过可变形Transformer显式捕捉视图间成对与高阶关联
- 在ModelNet40等数据集上达最新水平,适合3D识别研究者
基于视图的方法在3D形状理解中表现优异,但通常对视图间关系有强假设或间接学习多视图关联,限制了探索视图间关联的灵活性与任务有效性。为此,本文研究灵活组织与显式关联学习。提出将3D形状的不同视图整合为排列不变的集合(即视图集),消除刚性关系假设,促进视图间充分的信息交换与融合。在此基础上,设计轻量级Transformer模型VSFormer,显式捕获集合内所有元素的成对及高阶关联。同时,理论揭示视图集的笛卡尔积与注意力机制中的相关矩阵之间的自然对应关系,支撑模型设计。大量实验表明,VSFormer具有更强灵活性、高效推理能力与优越性能。尤其在ModelNet40、ScanObjectNN、RGBD等3D识别数据集上达到当前最优结果,并在SHREC'17检索基准上建立新纪录。
原文摘要 · Abstract (English)
View-based methods have demonstrated promising performance in 3D shape understanding. However, they tend to make strong assumptions about the relations between views or learn the multi-view correlations indirectly, which limits the flexibility of exploring inter-view correlations and the effectiveness of target tasks. To overcome the above problems, this paper investigates flexible organization and explicit correlation learning for multiple views. In particular, we propose to incorporate different views of a 3D shape into a permutation-invariant set, referred to as \emph{View Set}, which removes rigid relation assumptions and facilitates adequate information exchange and fusion among views. Based on that, we devise a nimble Transformer model, named \emph{VSFormer}, to explicitly capture pairwise and higher-order correlations of all elements in the set. Meanwhile, we theoretically reveal a natural correspondence between the Cartesian product of a view set and the correlation matrix in the attention mechanism, which supports our model design. Comprehensive experiments suggest that VSFormer has better flexibility, efficient inference efficiency and superior performance. Notably, VSFormer reaches state-of-the-art results on various 3d recognition datasets, including ModelNet40, ScanObjectNN and RGBD. It also establishes new records on the SHREC'17 retrieval benchmark. The code and datasets are available at \url{https://github.com/auniquesun/VSFormer}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。