arXiv:2608.26724cs.CV2026-08

提出几何感知的多视角异常检测框架,实现高效精准的缺陷定位。

GeoMAD: Geometry-Aware Multi-View Anomaly Detection via Deformable Fusion and Distributional Alignment

论文配图:GeoMAD: Geometry-Aware Multi-View Anomaly Detection via Deformable Fusion and Distributional Alignment
图 1 · 摘自论文原文
  • 通过可变形融合模块建立跨视角连续对应关系,无需标定或3D重建。
  • 在Real-IAD和MANTA-Tiny上达到先进性能,异常检测准确率达96.7%。
  • 适合工业质检场景,支持多类别、无监督异常检测任务。

多视角异常检测(MvAD)通过多个相机视角的互补观测来识别缺陷。核心挑战在于如何在保持可扩展性的同时实现充分的几何感知融合。现有方法通常处于两个极端:基于体素的融合虽有显式几何对齐但需昂贵的3D构建和类别特定假设;轻量级的块匹配融合虽高效,却依赖离散候选匹配且缺乏连续跨视角对应。本文提出GeoMAD,一种统一的多视角、多类别异常检测框架,同时解决几何对应不足与分布不一致问题。其提出的跨视角可变形融合模块(CDFM)直接在2D特征图上学习内容自适应、视图对特定的采样偏移,并结合图像全局参考采样构建多尺度窗口金字塔,实现无需相机标定、体素构建或类别特定3D监督的层级跨视角对应。进一步引入分布视角对齐(DVA),一种自监督跨视角正则化损失,将每视图的瓶颈分布对齐至实例级视图中心目标,强制全局一致性而无需像素级对应。CDFM与DVA协同,兼顾局部几何对应与全局分布一致性,在保留2D特征空间学习效率的同时,实现几何感知且分布一致的融合。在Real-IAD与MANTA-Tiny上的大量实验表明,GeoMAD在统一多视角异常检测中表现出色。

原文摘要 · Abstract (English)

Multi-view anomaly detection (MvAD) detects defects by exploiting complementary observations from multiple camera viewpoints. The central challenge is to fuse views with sufficient geometric awareness while remaining scalable to multi-class industrial settings. Existing methods typically fall into two extremes: voxel-based fusion provides explicit geometric alignment but requires costly 3D construction and class-specific assumptions, whereas lightweight patch-based fusion is efficient but relies on discrete candidate matching and lacks continuous cross-view correspondence. In this paper, we propose GeoMAD, a unified multi-view, multi-class AD framework that addresses both geometric correspondence deficiency and distributional inconsistency. Our \textit{Cross-view Deformable Fusion Module} (CDFM) learns content-adaptive, view-pair-specific sampling offsets directly on 2D feature maps and arranges them across a multi-scale window pyramid with image-global reference sampling, enabling hierarchical cross-view correspondence without camera calibration, voxel construction, or class-specific 3D supervision. We further introduce \textit{Distributional View Alignment} (DVA), a self-supervised cross-view regularization loss that aligns each view's bottleneck distribution against a per-instance view-centric target, enforcing global consistency without pixel-level correspondence. Together, CDFM and DVA bridge local geometric correspondence and global distributional consistency, providing geometry-aware and distribution-consistent fusion while preserving the efficiency of 2D feature-space learning. Extensive experiments on Real-IAD and MANTA-Tiny show that GeoMAD achieves strong detection and localization performance in unified MvAD.

异常检测多视角几何感知工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。