分两步训练提升声音事件定位与检测精度
A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection
- 先保持时间一致性,再分别训练事件检测与方向估计模型
- 在DCASE 2023数据集上显著提升分类与定位准确率
- 适合需要高精度空间音频感知的智能系统开发
声音事件定位与检测(SELD)在空间音频处理中至关重要,可识别声音事件并估计其三维方向。现有方法多采用单分支或双分支架构:单分支共享事件检测(SED)与到达方向(DoA)特征,导致优化冲突;双分支虽分离任务但限制信息交互。为此,本文提出一种两步学习框架:首先引入逐事件重排序格式,确保时间一致性,防止事件跨轨迹错配;其次分别训练SED与DoA网络,避免干扰,实现任务特异性特征学习;最后有效融合两者特征,增强空间与事件表征。在2023 DCASE挑战赛任务3数据集上的实验验证了该框架的有效性,成功克服单、双分支局限,显著提升事件分类与定位性能。
原文摘要 · Abstract (English)
Sound Event Localization and Detection (SELD) is crucial in spatial audio processing, enabling systems to detect sound events and estimate their 3D directions. Existing SELD methods use single- or dual-branch architectures: single-branch models share SED and DoA representations, causing optimization conflicts, while dual-branch models separate tasks but limit information exchange. To address this, we propose a two-step learning framework. First, we introduce a tracwise reordering format to maintain temporal consistency, preventing event reassignments across tracks. Next, we train SED and DoA networks to prevent interference and ensure task-specific feature learning. Finally, we effectively fuse DoA and SED features to enhance SELD performance with better spatial and event representation. Experiments on the 2023 DCASE challenge Task 3 dataset validate our framework, showing its ability to overcome single- and dual-branch limitations and improve event classification and localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。