arXiv:2605.10833cs.CVcs.AI2026-05

首个面向工业缺陷检测的多视角连续视频数据集及评测基准

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

论文配图:MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
图 1 · 摘自论文原文
  • 构建首个连续多视角工业视频数据集,支持多任务评估
  • 新模型VISTA在四项任务上平均得分提升至57.5,超越GPT-5.4
  • 适合工业视觉、视频理解与缺陷检测研究者使用

工业异常检测对制造质量控制至关重要,但现有数据集多聚焦静态图像或稀疏视角,难以反映真实工业场景中的连续检测过程。本文提出MMVIAD(多视角多任务工业异常检测),据知是首个用于工业异常检测与理解的连续多视角视频数据集,并建立多任务评测基准。该数据集包含以物体为中心的2秒检测片段,约120度相机运动,涵盖48类物体、14种环境和6种结构异常类型,支持异常检测、缺陷分类、物体分类及异常可见时间定位。在MMVIAD上的系统评估表明,当前商业与开源视频多模态大模型仍远低于人类表现,尤其在细粒度缺陷识别与时间定位方面。为提升可迁移的异常理解能力,本文进一步提出两阶段后训练流程:先通过感知结构微调(PS-SFT)建立感知结构化推理,再通过基于语义门控缺陷奖励与可视性感知时间奖励的VISTA-GRPO优化,得到最终模型VISTA。在未见数据集MMVIAD-Unseen上,VISTA将基线模型平均得分从45.0提升至57.5,超越GPT-5.4。代码已开源。

原文摘要 · Abstract (English)

Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which do not fully reflect continuous inspection processes in real industrial scenarios. We introduce MMVIAD (Multi-view Multi-task Video Industrial Anomaly Detection), to the best of our knowledge the first continuous multi-view video dataset for industrial anomaly detection and understanding, together with a benchmark for multi-task evaluation. MMVIAD contains object-centric 2-second inspection clips with approximately 120 degrees of camera motion, covering 48 object categories, 14 environments, and 6 structural anomaly types. It supports anomaly detection, defect classification, object classification, and anomaly visible-time localization. Systematic evaluations on MMVIAD show that current commercial and open-source video MLLMs remain far below human performance, especially for fine-grained defect recognition and temporal grounding. To improve transferable anomaly understanding, we further develop a two-stage post-training pipeline where PS-SFT (Perception-Structured Supervised Fine-Tuning) initializes perception-structured reasoning and VISTA-GRPO (Visibility-grounded Industrial Structured Temporal Anomaly Group Relative Policy Optimization) refines the model with semantic-gated defect reward and visibility-aware temporal reward, producing the final model VISTA. On MMVIAD-Unseen, VISTA improves the base model's average score across the four tasks from 45.0 to 57.5, surpassing GPT-5.4. Source code is available at https://github.com/Georgekeepmoving/MMVIAD.

工业检测视频理解多任务学习异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。