arXiv:2512.21650cs.LG2025-12

通过物理逻辑建模,提升多模态异常检测准确性

Physic-HM: Restoring Physical Generative Logic in Multimodal Anomaly Detection via Hierarchical Modulation

  • 用传感器引导音视频特征提取,建立过程到结果的单向生成关系
  • 在Weld-4M数据集上达到90.7%的I-AUROC,领先当前最优
  • 适合工业质检场景,尤其适用于焊接等复杂制造流程

多模态无监督异常检测对智能制造中的质量保障至关重要,尤其是在机器人焊接等复杂工艺中。现有方法常因忽视过程与结果间的单向物理生成逻辑,将过程模态(如实时视频、音频、传感器)与结果模态(如焊后图像)视为对称特征源,导致过程逻辑盲区。同时,高维视觉数据与低维传感器信号间的异质性使关键过程上下文被淹没。本文提出Physic-HM框架,显式引入物理归纳偏置,建模从过程到结果的依赖关系。核心创新包括:1)传感器引导的物理层级调制机制,利用低维传感器信号指导高维音视频特征提取;2)物理层级架构,强制实现单向生成映射以识别违反物理一致性的异常。在Weld-4M基准上的大量实验表明,Physic-HM达到90.7%的SOTA I-AUROC。代码将在论文录用后开源。

原文摘要 · Abstract (English)

Multimodal Unsupervised Anomaly Detection (UAD) is critical for quality assurance in smart manufacturing, particularly in complex processes like robotic welding. However, existing methods often suffer from process-logic blindness, treating process modalities (e.g., real-time video, audio, and sensors) and result modalities (e.g., post-weld images) as symmetric feature sources, thereby ignoring the inherent unidirectional physical generative logic. Furthermore, the heterogeneity gap between high-dimensional visual data and low-dimensional sensor signals frequently leads to critical process context being drowned out. In this paper, we propose Physic-HM, a multimodal UAD framework that explicitly incorporates physical inductive bias to model the process-to-result dependency. Specifically, our framework incorporates two key innovations: a Sensor-Guided PHM Modulation mechanism that utilizes low-dimensional sensor signals as context to guide high-dimensional audio-visual feature extraction, and a Physic-Hierarchical architecture that enforces a unidirectional generative mapping to identify anomalies that violate physical consistency. Extensive experiments on Weld-4M benchmark demonstrate that Physic-HM achieves a SOTA I-AUROC of 90.7%. The source code of Physic-HM will be released after the paper is accepted.

异常检测多模态工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。