arXiv:2607.18142cs.CVcs.AI2026-07中稿 · ECCV

无需训练,通过追踪物体状态变化来发现工业视频中的异常。

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

论文配图:O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
图 1 · 摘自论文原文
  • 基于物体中心的时空轨迹追踪与推理,模拟人工质检员思路。
  • 在三个工业数据集上超越主流VLM与传统方法,性能领先。
  • 无需领域知识或微调,可生成可解释的异常报告,适合工业质检场景。

工业视频异常检测(IVAD)旨在识别工业流程中的异常物体与事件,对现代制造与质量控制至关重要。现有基于视觉语言模型(VLM)的方法虽能检测通用场景下的开放域异常,但在具有复杂物体变换、严格物理规律和流程约束的工业场景中表现下降。为此,我们提出一种无需训练的代理式框架,不依赖特定领域知识,强调如人工质检员般关注物体状态演化。该方法追踪检测到物体在时空上的动态与内在变换,进而对物体级的时间状态轨迹进行推理,以识别真实帧中的异常物体。相比依赖正常片段重训或测试时注入领域知识的方法,本方案更具普适性。在三个IVAD数据集上的实验表明,我们的方法优于前沿VLM、代理框架及微调后的传统方法,并能提供关于异常过程与类型的可解释报告。

原文摘要 · Abstract (English)

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free of domain-specific knowledge, emphasizing object state evolution like humans inspectors. It is designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames. Our method overcomes limitations of prior approaches that rely on retraining on normal clips or injecting domain knowledge as context for test-time inference. Extensive experiments on three IVAD datasets demonstrate that our method outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine-tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.

工业视觉异常检测物体追踪可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。