arXiv:2501.10755cs.SDcs.LG2025-01中稿 · ICASSP2025被引 6

提出三种联合建模方法,实现声源三维定位与距离估计

An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation

  • 将方向与距离估计分步训练后融合,提升3D定位精度
  • 在DCASE 2024挑战赛中排名第一,验证方法有效性
  • 适合语音定位、智能音频系统等应用研究者参考

传统声音事件定位与检测(SELD)任务主要关注声音事件检测(SED)和到达方向(DOA)估计,但无法提供声源完整的空间信息。3D SELD任务通过引入源距离估计(SDE)弥补这一缺陷,实现完全的空间定位。本文提出三种解决方案:一种先独立训练再联合预测的新方法,将DOA与距离估计视为独立任务后融合;一种采用源笛卡尔坐标表示的双分支结构,实现DOA与距离的同步估计;以及一个三分支统一框架,联合建模SED、DOA与SDE。所提方法在DCASE 2024挑战赛任务3中排名第一,证明了联合建模在解决3D SELD任务中的有效性。相关代码未来将开源。

原文摘要 · Abstract (English)

In traditional sound event localization and detection (SELD) tasks, the focus is typically on sound event detection (SED) and direction-of-arrival (DOA) estimation, but they fall short of providing full spatial information about the sound source. The 3D SELD task addresses this limitation by integrating source distance estimation (SDE), allowing for complete spatial localization. We propose three approaches to tackle this challenge: a novel method with independent training and joint prediction, which firstly treats DOA and distance estimation as separate tasks and then combines them to solve 3D SELD; a dual-branch representation with source Cartesian coordinate used for simultaneous DOA and distance estimation; and a three-branch structure that jointly models SED, DOA, and SDE within a unified framework. Our proposed method ranked first in the DCASE 2024 Challenge Task 3, demonstrating the effectiveness of joint modeling for addressing the 3D SELD task. The relevant code for this paper will be open-sourced in the future.

声源定位三维感知多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。