arXiv:2605.11723cs.CVcs.AI2026-05被引 4

通过分层时空聚焦,提升视频异常检测的精准与可解释性。

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

论文配图:CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
图 1 · 摘自论文原文
  • 采用粗到精策略,先定位异常时间段,再精细定位空间区域。
  • 在细粒度异常基准上准确率提升25.7%,生成视频异常减少11.7%。
  • 首个带逐帧框标注和归因标签的大规模视频异常数据集。

本文提出一种基于视觉语言模型的粗到精异常奖励模型CaC。推理时,先全局扫描时间轴锚定异常时间段,再在局部区间内进行细粒度空间定位,最终通过结构化时空思维链推理得出稳健判断。为赋予模型该能力,我们构建了首个大规模生成视频异常数据集,包含逐帧边界框标注、时间异常窗口及细粒度归因标签。基于此数据集,设计三阶段渐进式训练范式:先通过单帧与多帧监督微调学习时空锚定,再利用双轮组相对策略优化(GRPO)进行强化学习优化。除传统准确率奖励外,引入时间和空间IoU奖励以监督中间定位过程,有效引导模型实现更可靠、可解释的时空推理。大量实验表明,CaC能稳定聚焦于细微异常,在细粒度异常基准上准确率提升25.7%;作为奖励信号使用时,可使生成视频异常减少11.7%,同时提升整体视频质量。

原文摘要 · Abstract (English)

In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiotemporal Chain-of-Thought reasoning. To equip the model with these capabilities, we construct the first large-scale generated video anomaly dataset with per-frame bounding-box annotations, temporal anomaly windows, and fine-grained attribution labels. Building on this dataset, we design a three-stage progressive training paradigm. The model initially learns spatial and temporal anchoring through single- and multi-frame supervised fine-tuning, and then is optimized by a reinforcement learning strategy based on two-turn Group Relative Policy Optimization (GRPO). Beyond conventional accuracy rewards, we introduce Temporal and Spatial IoU rewards to supervise the intermediate localization process, effectively guiding the model toward more grounded and interpretable spatiotemporal reasoning. Extensive experiments demonstrate that CaC can stably concentrate on subtle anomalies, achieving a 25.7% accuracy improvement on fine-grained anomaly benchmarks and, when used as a reward signal, CaC reduces generated-video anomalies by 11.7% while improving overall video quality.

视频异常奖励模型时空聚焦视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。