arXiv:2504.16102cs.CVcs.RO2025-04被引 1

针对音视频监控中怠速车辆检测难题,提出异构感知网络提升识别准确率。

HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues

  • 设计视觉金字塔与解耦检测头,应对多尺度与模态差异
  • 在公开数据集上相比基线模型提升9.42% mAP
  • 适合需要高精度车辆状态识别的智能交通场景

怠速车辆检测(IVD)通过监控视频与多通道音频,定位并分类拾取区域中车辆在最后一帧的状态(行驶、怠速或发动机关闭)。该任务面临三大挑战:(i)视觉线索与音频模式间的模态异质性;(ii)检测框尺度变化大,需多分辨率检测;(iii)因检测头耦合导致训练不稳定。现有端到端(E2E)模型采用简单CBAM双模态注意力,难以应对上述问题,常遗漏目标。本文提出HAVT-IVD,一种具备异质性感知能力的网络,包含视觉特征金字塔与解耦检测头。实验表明,该模型相较非联合基线提升mAP 7.66%,相较端到端基线提升9.42%。

原文摘要 · Abstract (English)

Idling vehicle detection (IVD) uses surveillance video and multichannel audio to localize and classify vehicles in the last frame as moving, idling, or engine-off in pick-up zones. IVD faces three challenges: (i) modality heterogeneity between visual cues and audio patterns; (ii) large box scale variation requiring multi-resolution detection; and (iii) training instability due to coupled detection heads. The previous end-to-end (E2E) model with simple CBAM-based bi-modal attention fails to handle these issues and often misses vehicles. We propose HAVT-IVD, a heterogeneity-aware network with a visual feature pyramid and decoupled heads. Experiments show HAVT-IVD improves mAP by 7.66 over the disjoint baseline and 9.42 over the E2E baseline.

音视频融合车辆检测智能交通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。