arXiv:2609.06978cs.CV2026-09

构建70万条细粒度异常数据,可控生成并验证图文一致性。

AnomalyCraft-700K: Component-Level Controllable and Verifiable Synthetic Anomalies for Fine-Grained Video Anomaly Understanding

论文配图:AnomalyCraft-700K: Component-Level Controllable and Verifiable Synthetic Anomalies for Fine-Grained Video Anomaly Understanding
图 1 · 摘自论文原文
  • 按语义组件分步生成异常事件,实现细粒度控制。
  • 含70万条标注、4万段视频,支持异常检测到推理全流程。
  • 适合研究细粒度视频异常理解与多模态对齐的学者。

视频异常理解(VAU)进展受限于真实异常视频难获取且内容不可控。现有合成方法多在类别或提示层控制,缺乏组件级验证与近似边界正常样本。为此,我们提出AnomalyCraft-700K,一个组件可控的合成异常数据集,包含超过4万段视频和70万条任务级文本标注。通过细粒度语义组件与三阶段渐进式流程,生成具有丰富细节、语义可控且时间结构化的异常事件,并为每类构建硬性正常样本,促使模型基于异常语义而非表面视觉特征进行判别。此外,以组件为单位校正生成中引入的图文不一致,确保六项任务(从异常检测到细粒度推理)的跨模态对齐可靠性。传统与多模态大模型协议下的评估表明,该数据集能有效提升从异常检测到细粒度理解的性能。

原文摘要 · Abstract (English)

Progress in video anomaly understanding (VAU) has long been limited by inherent deficiencies of real-world anomaly videos, which are hard to collect and offer little control over their content. Synthetic anomaly approaches partially alleviate data scarcity, yet their generation remains largely controlled at the category or prompt level. They also lack component-level verification of video-text consistency and provide insufficient hard normal samples near the normal-anomaly boundary. To address this, we present AnomalyCraft-700K, a component-controllable synthetic anomaly dataset for fine-grained VAU, containing over 40K videos and over 700K task-level textual annotations. From fine-grained semantic components and a progressive three-stage pipeline, we craft anomaly events that are richly detailed, semantically controlled, and temporally structured, and additionally construct per-category hard normal samples to prompt the model to discriminate based on anomaly semantics rather than surface visual cues. Moreover, using the components as verification units, AnomalyCraft-700K further performs component-wise correction of video-text discrepancies introduced during generation, providing reliable annotations with verified cross-modal alignment for six tasks that progress from anomaly detection, through anomaly retrieval and captioning, to fine-grained anomaly reasoning. Evaluations of widely used methods under both traditional and MLLM-based protocols demonstrate that AnomalyCraft-700K serves as an effective source of supervision, from anomaly detection to fine-grained anomaly understanding.

视频异常合成数据细粒度多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。