arXiv:2410.22271eess.AScs.AI2024-10被引 9

融合混响与视觉深度信息,提升声音事件定位与距离估计精度

Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation

  • 采用音视频融合的Conformer模型,结合ResNet50与预训练音频编码器
  • DOAE降低一半,F1值提升三倍以上,显著超越基线模型
  • 引入混响成分与深度图特征,适合多模态音频感知研究者参考

本文介绍我们提交至DCASE2024任务3的系统:基于音视频的声音事件定位与检测(含源距离估计)(赛道B)。主模型基于音视频Conformer,分别处理由ResNet50提取的视频特征和在SELDA上预训练的音频编码器生成的音频嵌入。该模型在STARSS23数据集开发集上表现优异,将DOAE降低约一半,F1值提升超过三倍。第二套系统通过时间集成融合AV-Conformer输出。随后引入直达声与混响成分(来自全向麦克风)及视频帧提取的深度图作为距离估计特征,新系统使RDE提升约3个百分点,但F1下降。分析表明可能因少数类声音样本训练不足导致检测能力下降。最终系统采用前三个模型预测的集成策略。未来可通过消融实验进一步优化系统与训练策略,实现性能持续提升。

原文摘要 · Abstract (English)

This report describes our systems submitted for the DCASE2024 Task 3 challenge: Audio and Audiovisual Sound Event Localization and Detection with Source Distance Estimation (Track B). Our main model is based on the audio-visual (AV) Conformer, which processes video and audio embeddings extracted with ResNet50 and with an audio encoder pre-trained on SELD, respectively. This model outperformed the audio-visual baseline of the development set of the STARSS23 dataset by a wide margin, halving its DOAE and improving the F1 by more than 3x. Our second system performs a temporal ensemble from the outputs of the AV-Conformer. We then extended the model with features for distance estimation, such as direct and reverberant signal components extracted from the omnidirectional audio channel, and depth maps extracted from the video frames. While the new system improved the RDE of our previous model by about 3 percentage points, it achieved a lower F1 score. This may be caused by sound classes that rarely appear in the training set and that the more complex system does not detect, as analysis can determine. To overcome this problem, our fourth and final system consists of an ensemble strategy combining the predictions of the other three. Many opportunities to refine the system and training strategy can be tested in future ablation experiments, and likely achieve incremental performance gains for this audio-visual task.

声音定位音视频融合距离估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。