arXiv:2601.02927cs.CVcs.AI2026-01中稿 · the 6th Workshop o…被引 1

用单个现成模型实现视频异常理解,高效又可解释。

PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding

  • 用文本锚点和提示优化实现粗略打分与精炼推理
  • 在标准数据集上表现接近先进方法,无需微调或外部模块
  • 适合实时应用,尤其对资源受限场景友好

视频异常理解(VAU)不仅定位异常,还需描述和推理其上下文。现有方法常依赖微调的多模态大语言模型(MLLM)或外部模块(如视频字幕生成器),导致标注成本高、训练复杂且推理开销大。本文提出PrismVAU,一种轻量高效的实时VAU系统,仅使用一个现成的MLLM完成异常评分、解释与提示优化。系统分两阶段:(1)基于与文本锚点的相似性计算帧级异常得分;(2)通过系统和用户提示,由MLLM进行上下文化推理。文本锚点与提示均通过弱监督自动提示工程(APE)框架优化。在标准VAD基准上的实验表明,PrismVAU在无需指令微调、帧级标注、外部模块或密集处理的前提下,实现了具有竞争力的检测性能和可解释的异常解释,为真实应用场景提供了高效实用的解决方案。

原文摘要 · Abstract (English)

Video Anomaly Understanding (VAU) extends traditional Video Anomaly Detection (VAD) by not only localizing anomalies but also describing and reasoning about their context. Existing VAU approaches often rely on fine-tuned multimodal large language models (MLLMs) or external modules such as video captioners, which introduce costly annotations, complex training pipelines, and high inference overhead. In this work, we introduce PrismVAU, a lightweight yet effective system for real-time VAU that leverages a single off-the-shelf MLLM for anomaly scoring, explanation, and prompt optimization. PrismVAU operates in two complementary stages: (1) a coarse anomaly scoring module that computes frame-level anomaly scores via similarity to textual anchors, and (2) an MLLM-based refinement module that contextualizes anomalies through system and user prompts. Both textual anchors and prompts are optimized with a weakly supervised Automatic Prompt Engineering (APE) framework. Extensive experiments on standard VAD benchmarks demonstrate that PrismVAU delivers competitive detection performance and interpretable anomaly explanations -- without relying on instruction tuning, frame-level annotations, and external modules or dense processing -- making it an efficient and practical solution for real-world applications.

视频异常多模态提示优化轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。