arXiv:2601.06204cs.CVcs.MA2026-01

用多智能体级联架构,让监控系统既快又懂语义。

Cascading multi-agent anomaly detection in surveillance systems via vision-language models and embedding-based classification

  • 分层处理:先快速筛选异常,再按需调用语义分析智能体。
  • 延迟降低3倍,图像质量保持在PSNR 38.3 dB、SSIM 0.965。
  • 适合需要实时+可解释性的智能监控场景。

动态视觉环境中的智能异常检测需兼顾实时性与语义可解释性。传统方法仅解决部分问题:基于重建的模型捕捉低层异常但缺乏上下文推理,目标检测器速度快但语义有限,大视觉语言系统虽可解释却计算成本过高。本文提出一种级联多智能体框架,融合上述互补范式,形成连贯可解释的架构。早期模块进行重建门控过滤与目标级评估,高层推理智能体仅在语义模糊事件时被调用。系统采用自适应升级阈值与发布-订阅通信机制,实现异步协调与跨异构硬件的可扩展部署。在大规模监控数据上的评估表明,该级联方案相比直接视觉语言推理,延迟降低三倍,同时保持高感知保真度(PSNR = 38.3 dB,SSIM = 0.965)和一致的语义标注。该框架超越传统检测流程,结合早退出效率、自适应多智能体推理与可解释异常归因,为可复现、低功耗的智能视觉监控提供基础。

原文摘要 · Abstract (English)

Intelligent anomaly detection in dynamic visual environments requires reconciling real-time performance with semantic interpretability. Conventional approaches address only fragments of this challenge. Reconstruction-based models capture low-level deviations without contextual reasoning, object detectors provide speed but limited semantics, and large vision-language systems deliver interpretability at prohibitive computational cost. This work introduces a cascading multi-agent framework that unifies these complementary paradigms into a coherent and interpretable architecture. Early modules perform reconstruction-gated filtering and object-level assessment, while higher-level reasoning agents are selectively invoked to interpret semantically ambiguous events. The system employs adaptive escalation thresholds and a publish-subscribe communication backbone, enabling asynchronous coordination and scalable deployment across heterogeneous hardware. Extensive evaluation on large-scale monitoring data demonstrates that the proposed cascade achieves a threefold reduction in latency compared to direct vision-language inference, while maintaining high perceptual fidelity (PSNR = 38.3 dB, SSIM = 0.965) and consistent semantic labeling. The framework advances beyond conventional detection pipelines by combining early-exit efficiency, adaptive multi-agent reasoning, and explainable anomaly attribution, establishing a reproducible and energy-efficient foundation for scalable intelligent visual monitoring.

异常检测多智能体视觉语言模型实时监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。