用事件流+视频流检测异常,提升复杂环境下的监控鲁棒性。
Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

- 融合事件流与可见光视频,利用事件相机高时序分辨率优势。
- 构建含63亿事件、37万帧的多模态数据集,支持真实场景研究。
- 自适应融合机制动态结合时空信息,适合工业级智能监控应用。
视频异常检测(VAD)在自动化监控中至关重要,但在光照变化、快速运动和复杂背景等挑战下,仅依赖可见光视频的方法仍显脆弱。为此,我们提出EVAD框架,联合使用生物启发式事件相机捕获的常规视频与事件流。事件传感器以高时间分辨率异步记录亮度变化,对运动模糊和极端光照具有鲁棒性,并提供与视频互补的运动显著特征。为支持多模态VAD研究,我们构建了一个大规模可见-事件基准数据集,包含63亿事件和376,368帧视频,覆盖多种光照水平、运动模式和背景复杂度,填补了事件驱动异常检测在真实性和可扩展性上的数据空白。基于该数据集,我们设计了一种对比多模态预训练框架,通过对齐事件流、可见视频与文本描述的语义嵌入,学习判别性事件表征。进一步引入自适应融合模块,动态整合事件流的时间线索与视频的空间语义,增强对环境扰动的鲁棒性。在多个基准及所提出的TJUTCM Pha数据集上的实验表明,EVAD持续优于现有方法,验证了事件感知在真实世界场景中对VAD的有效性。
原文摘要 · Abstract (English)
Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos. To address these limitations, we propose EVAD, an event enhanced VAD framework that jointly exploits conventional video and event streams captured by bio inspired event cameras. Event sensors asynchronously capture brightness changes with high temporal resolution, offering robustness to motion blur and extreme lighting, and providing motion salient cues complementary to video based visual information. To support multi modal VAD research, we construct a large scale visible event benchmark comprising 6.3 billion events and 376,368 video frames collected under diverse illumination levels, motion patterns, and background complexities, filling the gap of realistic and scalable datasets for event based anomaly detection. Building upon this dataset, we design a contrastive multi modal pretraining framework to learn discriminative event representations by aligning semantic embeddings across event streams, visible videos, and textual descriptions. An adaptive fusion module then dynamically integrates event based temporal cues with video based spatial semantics, improving robustness to environmental disturbances. Experiments on benchmarks and the proposed TJUTCM Pha dataset demonstrate that E VAD consistently outperforms methods, validating the effectiveness of event-based sensing for VAD in real world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。