跨视图跨模态映射,提升3D异常检测精度
Modulate-and-Map: Crossmodal Feature Mapping with Cross-View Modulation for 3D Anomaly Detection
- 通过特征调制实现多视角多模态特征映射
- 在SiM3D上超越现有方法,显著提升异常检测性能
- 适合工业场景的高分辨率3D数据异常分析
我们提出ModMap,一种原生多视图多模态的3D异常检测与分割框架。不同于以往独立处理各视图的方法,本方法借鉴跨模态特征映射思想,学习在不同模态和视图间映射特征,并通过逐特征调制显式建模视图依赖关系。引入跨视图训练策略,利用所有可能的视图组合,实现多视图集成与聚合的有效异常评分。为处理高分辨率3D数据,我们训练并公开发布了一个专用于工业数据集的基础深度编码器。在最新基准SiM3D上的实验表明,该方法在多视图多模态设置下取得当前最优性能,显著优于先前方法。
原文摘要 · Abstract (English)
We present ModMap, a natively multiview and multimodal framework for 3D anomaly detection and segmentation. Unlike existing methods that process views independently, our method draws inspiration from the crossmodal feature mapping paradigm to learn to map features across both modalities and views, while explicitly modelling view-dependent relationships through feature-wise modulation. We introduce a cross-view training strategy that leverages all possible view combinations, enabling effective anomaly scoring through multiview ensembling and aggregation. To process high-resolution 3D data, we train and publicly release a foundational depth encoder tailored to industrial datasets. Experiments on SiM3D, a recent benchmark that introduces the first multiview and multimodal setup for 3D anomaly detection and segmentation, demonstrate that ModMap attains state-of-the-art performance by surpassing previous methods by wide margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。