提出新方法提升复杂声音环境下的声事件定位精度。
Location-Oriented Sound Event Localization and Detection with Spatial Mapping and Regression Localization
- 将三维空间映射到二维平面,用回归损失优化定位结果。
- 在STARSS22/23数据集上超越现有方法,多音重叠场景表现更优。
- 适合需要高鲁棒性声事件定位的智能音频系统开发者。
声事件定位与检测(SELD)将声事件检测(SED)与到达方向(DOA)结合。现有基于事件的多轨方法在多重音环境下因轨道数限制而泛化能力不足。为此,本文提出面向位置的声事件定位与检测方法(SMRL-SELD),将三维空间分割并映射至二维平面,设计新的回归定位损失,使模型输出更贴近真实事件位置。该方法以位置为导向,使模型能根据方向学习事件特征,从而不受重叠事件数量限制地处理多音环境。在STARSS22和STARSS23数据集上的实验表明,SMRL-SELD在整体性能及多音场景下均优于现有方法。
原文摘要 · Abstract (English)
Sound Event Localization and Detection (SELD) combines the Sound Event Detection (SED) with the corresponding Direction Of Arrival (DOA). Recently, adopted event oriented multi-track methods affect the generality in polyphonic environments due to the limitation of the number of tracks. To enhance the generality in polyphonic environments, we propose Spatial Mapping and Regression Localization for SELD (SMRL-SELD). SMRL-SELD segments the 3D spatial space, mapping it to a 2D plane, and a new regression localization loss is proposed to help the results converge toward the location of the corresponding event. SMRL-SELD is location-oriented, allowing the model to learn event features based on orientation. Thus, the method enables the model to process polyphonic sounds regardless of the number of overlapping events. We conducted experiments on STARSS23 and STARSS22 datasets and our proposed SMRL-SELD outperforms the existing SELD methods in overall evaluation and polyphony environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。