提出端到端音视频联合检测怠速车辆的新方法
Joint Audio-Visual Idling Vehicle Detection with Streamlined Input Dependencies
- 用双向注意力融合音视频特征,实现端到端检测
- 在七倍于以往的数据集上达到可部署的准确率
- 适合自动驾驶车载摄像头系统直接应用
怠速车辆检测有助于监控和减少不必要的怠速,可集成到实时系统中以缓解污染与有害排放。此前方法[13]为非端到端模型,需人工点击指定输入区域,导致部署易出错甚至不可行。本文提出一种端到端的音视频联合怠速车辆检测任务,用于识别车辆在行驶、怠速、发动机关闭三种状态。不同于音视频协同定位等特征共现任务,本任务依赖双模态互补信息,单模态无法独立确定标签。为此,提出AVIVD-Net,通过双向注意力机制融合音视频特征,构建联合特征空间,简化输入流程,降低部署复杂度。此外,构建了规模达以往七倍的AVIVD数据集,提供更丰富的标注样本。模型性能与先前方法相当,具备自动化部署潜力。进一步在公开的特征共现数据集MAVD[23]上评估,验证其在自动驾驶车载摄像头场景下的扩展性。
原文摘要 · Abstract (English)
Idling vehicle detection (IVD) can be helpful in monitoring and reducing unnecessary idling and can be integrated into real-time systems to address the resulting pollution and harmful products. The previous approach [13], a non-end-to-end model, requires extra user clicks to specify a part of the input, making system deployment more error-prone or even not feasible. In contrast, we introduce an end-to-end joint audio-visual IVD task designed to detect vehicles visually under three states: moving, idling and engine off. Unlike feature co-occurrence task such as audio-visual vehicle tracking, our IVD task addresses complementary features, where labels cannot be determined by a single modality alone. To this end, we propose AVIVD-Net, a novel network that integrates audio and visual features through a bidirectional attention mechanism. AVIVD-Net streamlines the input process by learning a joint feature space, reducing the deployment complexity of previous methods. Additionally, we introduce the AVIVD dataset, which is seven times larger than previous datasets, offering significantly more annotated samples to study the IVD problem. Our model achieves performance comparable to prior approaches, making it suitable for automated deployment. Furthermore, by evaluating AVIVDNet on the feature co-occurrence public dataset MAVD [23], we demonstrate its potential for extension to self-driving vehicle video-camera setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。