首个基于物理规律的工业异常检测数据集,助力机器像人一样推理物体异常。
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
- 构建真实机器人操作场景下的动态视频数据集,融合物理规律与视觉内容。
- 涵盖6400+视频、22类物体、47种异常,需结合物理知识进行视觉推理。
- 引入可解释性评估指标,适合研究物理感知与可解释性模型的学者。
人类通过感知、交互和基于物体条件的物理知识来识别现实世界中的物体异常。工业异常检测(IAD)的长期目标是让机器自主具备这一能力。然而,当前IAD算法主要在静态、语义简单的数据集上开发和测试,与需要物理理解与推理的真实场景存在差距。为此,我们提出了首个大规模、真实世界、以物理为基础的视频数据集——物理异常检测(Phys-AD)。该数据集由真实机械臂和电机采集,包含超过6400个视频,覆盖22类真实物体,涉及多种动态交互,共包含47种异常类型。在Phys-AD中进行异常检测需要结合物理知识与视频内容进行视觉推理。我们在三种设置下对前沿异常检测方法进行了基准测试:无监督、弱监督和视频理解任务,揭示了现有方法在处理物理相关异常时的局限性。此外,我们提出了物理异常解释(PAEval)评估指标,用于衡量视觉-语言基础模型不仅能够检测异常,还能准确解释其物理成因的能力。项目主页:https://guyao2023.github.io/Phys-AD/
原文摘要 · Abstract (English)
Humans detect real-world object anomalies by perceiving, interacting, and reasoning based on object-conditioned physical knowledge. The long-term goal of Industrial Anomaly Detection (IAD) is to enable machines to autonomously replicate this skill. However, current IAD algorithms are largely developed and tested on static, semantically simple datasets, which diverge from real-world scenarios where physical understanding and reasoning are essential. To bridge this gap, we introduce the Physics Anomaly Detection (Phys-AD) dataset, the first large-scale, real-world, physics-grounded video dataset for industrial anomaly detection. Collected using a real robot arm and motor, Phys-AD provides a diverse set of dynamic, semantically rich scenarios. The dataset includes more than 6400 videos across 22 real-world object categories, interacting with robot arms and motors, and exhibits 47 types of anomalies. Anomaly detection in Phys-AD requires visual reasoning, combining both physical knowledge and video content to determine object abnormality. We benchmark state-of-the-art anomaly detection methods under three settings: unsupervised AD, weakly-supervised AD, and video-understanding AD, highlighting their limitations in handling physics-grounded anomalies. Additionally, we introduce the Physics Anomaly Explanation (PAEval) metric, designed to assess the ability of visual-language foundation models to not only detect anomalies but also provide accurate explanations for their underlying physical causes. Our project is available at https://guyao2023.github.io/Phys-AD/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。