arXiv:2604.23724cs.CVcs.AI2026-04

用贝叶斯推理聚焦远距离车辆异常,提升交通监控效率与准确性。

Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference

  • 基于贝叶斯推断动态建模正常车流运动分布,实时更新异常判断边界。
  • 仅对触发区域进行视觉语言模型推理,减少90%以上无效计算。
  • 适用于复杂高速场景,兼顾实时性与语义理解能力,适合智能交通系统部署。

高速公路视频异常检测对交通安全至关重要,但远距离车辆的细微异常运动在多样场景下仍具挑战。视觉-语言模型(VLM)具备强大语义推理能力,但全帧处理会稀释远端目标证据并带来巨大计算开销。为此,我们提出VIBES异步框架,利用贝叶斯推断引导聚焦式VLM推理。具体地,一个在线运动引导贝叶斯推断模块持续从车辆轨迹中估计上下文相关的正常运动分布,并更新其概率边界。偏离该边界即产生异步触发信号,定位时间与空间上的候选异常。系统不处理连续全帧视频,而仅对触发关联的帧和局部视觉区域进行推理,有效减少无关内容与冗余计算。大量实验表明,VIBES在多种高速路条件下显著提升远距离异常检测性能与语义解释能力,同时实现实时处理效率。

原文摘要 · Abstract (English)

Expressway video anomaly detection is important for traffic safety, but remains challenging across diverse scenes, particularly for far-field vehicles with subtle abnormal motion. Vision-Language Models (VLMs) provide strong semantic reasoning capabilities, yet processing full frames can dilute evidence from distant targets and introduce substantial computational overhead. To address these challenges, we propose VIBES, an asynchronous framework that uses Bayesian inference to guide focused VLM reasoning. Specifically, an online kinematics-guided Bayesian inference module continuously estimates a context-dependent normal-motion distribution from vehicle trajectories and updates its probabilistic boundaries. Deviations from these boundaries produce asynchronous triggers that localize candidate anomalies in time and space. Instead of processing continuous full-frame video, the VLM reasons only over selected frames and localized visual regions associated with the triggers, reducing irrelevant visual content and unnecessary inference. Extensive experiments show that VIBES improves far-field anomaly detection and semantic interpretation while achieving real-time processing efficiency across diverse expressway conditions.

异常检测视觉语言模型贝叶斯推理交通监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。