让大模型精准定位视频异常,提升对稀疏异常的识别能力。
Aligning Effective Tokens with Video Anomaly in Large Language Models
- 通过空间与时间有效令牌对齐,增强模型对异常事件的感知。
- 在XD-Violence等数据集上超越现有方法,准确率显著提升。
- 专为异常检测设计数据集与评测基准,适合视频安全研究者。
理解视频中的异常事件是一项关键且具挑战性的任务,在众多应用中备受关注。尽管当前多模态大语言模型(MLLMs)能够分析一般视频,但因异常事件在时空上具有稀疏性,冗余信息常导致性能不佳。为此,我们提出VA-GPT,一种新型MLLM,用于总结和定位各类视频中的异常事件。通过引入两个核心模块——空间有效令牌选择(SETS)和时间有效令牌生成(TETG),该方法高效对齐视觉编码器与大语言模型之间的有效令牌,从而更准确地捕捉和分析异常事件相关的时空信息,实现更优响应与交互。此外,我们构建了一个面向指令跟随的视频异常感知微调数据集,并基于XD-Violence数据集建立跨域评估基准。实验表明,所提方法在多个基准上优于现有最先进方法。
原文摘要 · Abstract (English)
Understanding abnormal events in videos is a vital and challenging task that has garnered significant attention in a wide range of applications. Although current video understanding Multi-modal Large Language Models (MLLMs) are capable of analyzing general videos, they often struggle to handle anomalies due to the spatial and temporal sparsity of abnormal events, where the redundant information always leads to suboptimal outcomes. To address these challenges, exploiting the representation and generalization capabilities of Vison Language Models (VLMs) and Large Language Models (LLMs), we propose VA-GPT, a novel MLLM designed for summarizing and localizing abnormal events in various videos. Our approach efficiently aligns effective tokens between visual encoders and LLMs through two key proposed modules: Spatial Effective Token Selection (SETS) and Temporal Effective Token Generation (TETG). These modules enable our model to effectively capture and analyze both spatial and temporal information associated with abnormal events, resulting in more accurate responses and interactions. Furthermore, we construct an instruction-following dataset specifically for fine-tuning video-anomaly-aware MLLMs, and introduce a cross-domain evaluation benchmark based on XD-Violence dataset. Our proposed method outperforms existing state-of-the-art methods on various benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。