针对音视频监控中怠速车辆检测难题,提出异构感知网络提升识别准确率。
HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues
- 设计视觉金字塔与解耦检测头,应对多尺度与模态差异
- 在公开数据集上相比基线模型提升9.42% mAP
- 适合需要高精度车辆状态识别的智能交通场景
怠速车辆检测(IVD)通过监控视频与多通道音频,定位并分类拾取区域中车辆在最后一帧的状态(行驶、怠速或发动机关闭)。该任务面临三大挑战:(i)视觉线索与音频模式间的模态异质性;(ii)检测框尺度变化大,需多分辨率检测;(iii)因检测头耦合导致训练不稳定。现有端到端(E2E)模型采用简单CBAM双模态注意力,难以应对上述问题,常遗漏目标。本文提出HAVT-IVD,一种具备异质性感知能力的网络,包含视觉特征金字塔与解耦检测头。实验表明,该模型相较非联合基线提升mAP 7.66%,相较端到端基线提升9.42%。
原文摘要 · Abstract (English)
Idling vehicle detection (IVD) uses surveillance video and multichannel audio to localize and classify vehicles in the last frame as moving, idling, or engine-off in pick-up zones. IVD faces three challenges: (i) modality heterogeneity between visual cues and audio patterns; (ii) large box scale variation requiring multi-resolution detection; and (iii) training instability due to coupled detection heads. The previous end-to-end (E2E) model with simple CBAM-based bi-modal attention fails to handle these issues and often misses vehicles. We propose HAVT-IVD, a heterogeneity-aware network with a visual feature pyramid and decoupled heads. Experiments show HAVT-IVD improves mAP by 7.66 over the disjoint baseline and 9.42 over the E2E baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。