arXiv:2608.03779cs.CV2026-08

多智能体协作探查验证,让视频异常理解更全面可信。

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding

论文配图:AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding
图 1 · 摘自论文原文
  • 分四个专业智能体,分别负责规则构建、搜索规划、观察与决策
  • 在三个数据集上超越零样本和强化学习基线,准确率提升显著
  • 无需训练,适合需要可解释性的异常检测场景

视频异常理解(VAU)旨在全面解析视频中的异常事件,要求模型不仅能识别异常,还需发现支持证据并解释根本原因,而不仅是简单检测。现有方法常依赖特定训练或有限观察,限制了泛化能力与证据覆盖范围。尽管单智能体方法支持自适应视频观察,但仍将探索、观察与决策整合于统一推理流程中,缺乏角色分工与结构化证据协调。为此,我们提出AgenticVAU——一种免训练的多智能体框架,将VAU视为探查-验证过程:先发现潜在异常,再通过针对性观察验证。该框架引入四类专用智能体,分别负责视觉规则构建、搜索规划、视频观察与最终决策,并通过共享的锚点注册表(anchor registry)——即绑定每次观察的共享证据记忆——进行通信。在该框架引导下,AgenticVAU持续交织宽时域探索、密集局部验证与跨区间对比,直至收集充分证据。我们在ECVA、UCF-Crime及VAU-Bench的MSAD子集上进行了大量实验,结果表明,AgenticVAU优于零样本推理与基于强化学习的基线方法,验证了多智能体协作在视频异常理解中的价值。

原文摘要 · Abstract (English)

Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences, discover their supporting evidence, and explain the underlying causes beyond simple anomaly detection. Existing VAU methods often rely on specialized training or limited observations, restricting generalization or evidence coverage. Although single-agent alternatives support adaptive video observation, they still integrate exploration, observation, and decision-making within a unified reasoning process, offering limited role specialization and structured evidence coordination. To address these limitations, we present AgenticVAU, a training-free multi-agent framework that casts VAU as an explore--verify process, where the system first discovers potential anomalies and then verifies them through targeted observations. To achieve this, four specialized agents are introduced to handle visual-rule construction, search planning, video observation, and final decision, respectively. These agents communicate through an anchor registry, a shared evidence memory that binds each observation. Guided by this agent framework, AgenticVAU interleaves broad temporal exploration, dense local verification, and cross-interval comparison until sufficient evidence is collected. We conduct extensive experiments on the ECVA, UCF-Crime, and MSAD subsets of VAU-Bench, the results show that AgenticVAU outperforms zero-shot inference and reinforcement learning-based baselines, demonstrating the value of multi-agent collaboration for video anomaly understanding.

视频异常多智能体可解释性推理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。