用生成式方法构建高复杂度、平衡的视频异常检测基准
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
- 全自动生成场景与异常事件,精准控制叙事时序
- 产出41秒连贯长视频,涵盖多样异常类型与动态事件
- 适合研究动态异常理解与因果推理的学者使用
自动检测视频中的异常事件对现代自主系统至关重要,但现有视频异常检测(VAD)基准在场景多样性、异常分布均衡性及时间复杂性方面不足,难以真实评估模型性能。与此同时,社区正转向视频异常理解(VAU),需更深层的语义与因果推理,却因标注成本过高而难于评测。本文提出Pistachio,一个完全通过可控生成式流程构建的VAD/VAU基准。利用最新视频生成模型,Pistachio可精确控制场景、异常类型与时间叙事,有效消除互联网数据集的偏差与局限。其流水线融合场景条件异常分配、多步故事生成与时间一致性长视频合成策略,实现仅需少量人工干预即可生成41秒连贯视频。大量实验验证了Pistachio的规模、多样性和复杂性,揭示了现有方法的新挑战,推动未来对动态与多重事件异常理解的研究。
原文摘要 · Abstract (English)
Automatically detecting abnormal events in videos is crucial for modern autonomous systems, yet existing Video Anomaly Detection (VAD) benchmarks lack the scene diversity, balanced anomaly coverage, and temporal complexity needed to reliably assess real-world performance. Meanwhile, the community is increasingly moving toward Video Anomaly Understanding (VAU), which requires deeper semantic and causal reasoning but remains difficult to benchmark due to the heavy manual annotation effort it demands. In this paper, we introduce Pistachio, a new VAD/VAU benchmark constructed entirely through a controlled, generation-based pipeline. By leveraging recent advances in video generation models, Pistachio provides precise control over scenes, anomaly types, and temporal narratives, effectively eliminating the biases and limitations of Internet-collected datasets. Our pipeline integrates scene-conditioned anomaly assignment, multi-step storyline generation, and a temporally consistent long-form synthesis strategy that produces coherent 41-second videos with minimal human intervention. Extensive experiments demonstrate the scale, diversity, and complexity of Pistachio, revealing new challenges for existing methods and motivating future research on dynamic and multi-event anomaly understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。