提出多阶段视频注意力网络,实现声音定位、分类与距离估计的3D声事件检测。
MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation
- 采用多阶段音频特征捕捉视频中声源空间信息。
- 通过笛卡尔坐标联合输出方向与距离,提升3D定位精度。
- 适用于需要精准声源位置感知的智能监控与机器人听觉系统。
基于音视频的三维声事件定位与检测(3D SELD)不仅需识别声音类别和到达方向(DOA),还需预测声源距离,以提供完整的声源位置信息。本文提出多阶段视频注意力网络(MVANet)用于音视频3D SELD。通过多阶段音频特征自适应提取视频中声源的空间信息,并设计一种新输出表示,将声源的到达方向与距离结合,计算真实笛卡尔坐标,以应对DCASE 2024挑战赛中新引入的声源距离估计(SDE)任务。同时采用多种有效数据增强与预训练方法。在STARSS23数据集上的实验结果表明,所提方法在不使用模型集成的情况下,优于该挑战赛音视频3D SELD任务的最先进方法。代码未来将公开。
原文摘要 · Abstract (English)
Sound event localization and detection with source distance estimation (3D SELD) involves not only identifying the sound category and its direction-of-arrival (DOA) but also predicting the source's distance, aiming to provide full information about the sound position. This paper proposes a multi-stage video attention network (MVANet) for audio-visual (AV) 3D SELD. Multi-stage audio features are used to adaptively capture the spatial information of sound sources in videos. We propose a novel output representation that combines the DOA with distance of sound sources by calculating the real Cartesian coordinates to address the newly introduced source distance estimation (SDE) task in the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 Challenge. We also employ a variety of effective data augmentation and pre-training methods. Experimental results on the STARSS23 dataset have proven the effectiveness of our proposed MVANet. By integrating the aforementioned techniques, our system outperforms the top-ranked method we used in the AV 3D SELD task of the DCASE 2024 Challenge without model ensemble. The code will be made publicly available in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。