让视频异常检测模型学会自我反思,提升判断准确性。
SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning
- 引入自省式思维链数据集,支持模型自我反思与修正。
- 在多个基准上实现异常定位精度和推理质量显著提升。
- 适合需要高可靠性视频分析的场景,如安防监控。
多模态大语言模型(MLLM)在推理能力上取得显著进展,并在视频异常理解(VAU)任务中展现出良好效果。然而,现有基于MLLM的方法仍主要停留在对异常的表面描述,缺乏对异常行为的深层推理,如显式的自我反思与自我修正。为此,我们提出自省增强型视频异常理解框架SRVAU-R1,该框架在MLLM推理中融入自省机制。具体而言,SRVAU-R1构建了首个面向VAU的自省式思维链数据集,提供包含初始推理、自我反思与修正推理的结构化监督信号。基于此,设计了一种新的自省感知学习范式,结合监督微调与强化微调,以提升多模态推理性能。在多个视频异常检测基准上的大量实验表明,SRVAU-R1持续优于现有方法,在时间异常定位准确率和推理质量方面均取得显著提升。
原文摘要 · Abstract (English)
Multi-modal large language models (MLLMs) have demonstrated significant progress in reasoning capabilities and shown promising effectiveness in video anomaly understanding (VAU) tasks. However, existing MLLM-based approaches remain largely focused on surface-level descriptions of anomalies, lacking deep reasoning over abnormal behaviors like explicit self-reflection and self-correction. To address that, we propose Self-Reflection-Enhanced Reasoning for Video Anomaly Understanding (SRVAU-R1), a reflection-aware learning framework that incorporates reflection in MLLM reasoning. Specifically, SRVAU-R1 introduces the first reflection-oriented Chain-of-Thought dataset tailored for VAU, providing structured supervision with initial reasoning, self-reflection, and revised reasoning. Based on that, it includes a novel reflection-aware learning paradigm with supervised fine-tuning and reinforcement fine-tuning to enhance multi-modal reasoning for VAU. Extensive experiments on multiple video anomaly benchmarks demonstrate that SRVAU-R1 consistently outperforms existing methods, achieving significant improvements in both temporal anomaly localization accuracy and reasoning quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。