提升立体声声音事件定位,用感知特征和专用增强方法。
Improving Stereo 3D Sound Event Localization and Detection: Perceptual Features, Stereo-specific Data Augmentation, and Distance Normalization
- 使用感知驱动的输入特征,增强事件检测与定位能力。
- 引入立体声专用数据增强,提升模型在真实场景下的泛化性。
- 适合音频定位、音视频融合研究者参考。
本技术报告介绍了我们参与 DCASE 2025 挑战赛任务3:常规视频内容中的立体声声音事件定位与检测(SELD)的方案。本文聚焦音频单模态任务,提出三项关键贡献:首先,设计了基于感知动机的输入特征,显著提升事件检测、声源定位及距离估计性能;其次,针对立体声音频特性,提出通道互换、时频掩蔽等专用数据增强策略,并引入尚未在 SELD 中应用的 FilterAugment 技术;最后,在训练中采用距离归一化方法,稳定回归目标。在立体声 STARSS23 数据集上的实验表明,所有 SELD 评价指标均实现持续提升。代码已公开于 https://github.com/itsjunwei/NTU_SNTL_Task3。
原文摘要 · Abstract (English)
This technical report presents our submission to Task 3 of the DCASE 2025 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. We address the audio-only task in this report and introduce several key contributions. First, we design perceptually-motivated input features that improve event detection, sound source localization, and distance estimation. Second, we adapt augmentation strategies specifically for the intricacies of stereo audio, including channel swapping and time-frequency masking. We also incorporate the recently proposed FilterAugment technique that has yet to be explored for SELD work. Lastly, we apply a distance normalization approach during training to stabilize regression targets. Experiments on the stereo STARSS23 dataset demonstrate consistent performance gains across all SELD metrics. Code to replicate our work is available in this repository: https://github.com/itsjunwei/NTU_SNTL_Task3
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。