arXiv:2607.00654cs.CV2026-07中稿 · ICML

无需参数更新,用语言描述异常经验提升视频异常检测效果

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

论文配图:Linguistic Relative Policy Optimization for Video Anomaly Reasoning
图 1 · 摘自论文原文
  • 通过多推理路径提炼相对语义优势,生成语言化异常先验
  • 在无调参条件下,在三个数据集上超越现有最优方法
  • 适合希望减少人工标注和规则设计的异常检测研究者

基于多模态大语言模型的视频异常检测展现出巨大潜力,但多数方法仍依赖大规模标注或专家设计的先验知识,限制了其在极低人工干预下获取异常知识的能力。为此,我们提出语言相对策略优化(LRPO),从多个推理轨迹中提炼群体相对语义优势,构建语言表达的异常经验先验,并通过注入上下文引导模型输出分布,无需任何参数更新。LRPO构建两种互补的经验表征:通用经验捕捉跨场景可迁移的异常偏好,场景经验建模上下文相关的异常规则以实现精准优化。为进一步提升学习到的经验,引入异常对齐奖励,引导轨迹优化匹配人类风险偏好并强化时序相关推理。在XD-Violence、UCF-Crime和UBnormal上的大量实验表明,LRPO在无需调参设置下显著优于现有最先进方法。

原文摘要 · Abstract (English)

Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale annotations or expert-designed priors, limiting their ability to acquire anomaly knowledge with as little human intervention as possible. To address this, we propose Linguistic Relative Policy Optimization (LRPO), which distills group-relative semantic advantages from multiple reasoning trajectories into a linguistically expressed anomaly experience prior, and adapts the model by injecting this prior into the context to steer its output distribution without any parameter updates. LRPO builds two complementary experience representations: general experience captures transferable anomaly preferences across scenarios, while scenario experience models context-dependent anomaly rules for targeted refinement. To further improve the learned experience, we introduce an anomaly alignment reward that guides trajectory optimization to match human risk preferences and reinforce temporally grounded reasoning. Extensive experiments on XD-Violence, UCF-Crime, and UBnormal demonstrate that LRPO significantly outperforms existing state-of-the-art methods under tuning-free settings.

视频异常检测大语言模型无监督学习经验先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。