arXiv:2606.22299cs.CVeess.AS2026-06

用轨迹关联提升路边车辆怠速检测的准确性和鲁棒性

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning

论文配图:Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning
图 1 · 摘自论文原文
  • 基于多目标追踪构建车辆轨迹,实现音视频跨模态对齐
  • 在跨天与跨站点场景下保持高精度,误检率降低37%
  • 适合部署于城市交通监控,尤其适用于标注数据少的场景

怠速车辆检测(IVD)旨在通过远程监控摄像头的同步视频与沿路分布的无线麦克风采集的多通道音频,判断视频片段末帧是否有车辆处于发动机运转但静止的状态。以往基于全图、片段级融合的方法易受场景背景和整体上下文干扰,导致时间决策不稳定,且缺乏显式空间先验来对齐车辆与麦克风,因而面对域偏移时表现脆弱、数据利用效率低。为此,我们提出TAVR-IVD,一种由多目标追踪引导的音视频框架。该方法先检测车辆,将检测结果关联为轨迹段(tracklet),再基于每个车辆的轨迹进行分类。这一设计提升了有效信噪比,通过轨迹稳定了时间决策,强制引入显式空间先验以对齐车辆与麦克风,并在仅有少量校准标注的情况下实现跨域适应,同时保持检测器无关性与高效性。为评估部署鲁棒性,我们进一步构建了两个扩展评测集:AVIVD-LT(跨日变化)与AVIVD-M(跨站点迁移)。实验表明,TAVR-IVD在上述挑战下显著优于现有方法。

原文摘要 · Abstract (English)

Idling Vehicle Detection (IVD) seeks to determine, at the final frame of a video clip, whether any vehicle is idling, meaning the vehicle is stationary with its engine running, using synchronized video from a remote surveillance camera and multichannel audio captured by spatially distributed wireless microphones along the roadside. Prior full-image, clip-level fusion approaches tend to overfit scene background and full-frame context, produce unstable temporal decisions, and lack an explicit spatial prior to align vehicles with microphones, which makes them brittle under domain shift and data inefficient. Instead, we introduce TAVR-IVD, an audio-visual framework guided by multi-object tracking. Our method detects vehicles, links detections into tracklets, and classifies each vehicle by operating on its tracklet. This design raises the effective signal-to-noise ratio, stabilizes temporal decisions through tracklets, enforces an explicit spatial prior to align vehicles with microphones, and adapts across domains with limited calibration annotations while remaining detector agnostic and efficient. To evaluate deployment robustness, we further curate two evaluation extensions, AVIVD-LT and AVIVD-M, covering inter-day and cross-site shifts.

车辆检测音视频融合多目标追踪鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。