arXiv:2506.02773eess.AScs.SD2025-06中稿 · and to appear at I…被引 6

AuralNet可精准定位重叠声源的三维方向,无需预知声源数量。

AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers

  • 分阶段处理:粗分类+细回归,支持灵活空间分辨率
  • 在混响噪声环境下定位准确率显著优于现有方法
  • 适合语音增强、智能听觉系统等需要精确定位的应用

我们提出AuralNet,一种新型的3D多声源双耳声音源定位方法,可在不预先知道声源数量的情况下,同时定位声源的方位角和仰角。AuralNet采用门控的粗粒度到细粒度架构,结合粗分类阶段与细粒度回归阶段,通过扇区划分实现灵活的空间分辨率。模型引入多头自注意力机制,以捕捉双耳信号中的空间线索,在混响噪声环境中表现出更强的鲁棒性。设计了一种掩码多任务损失函数,联合优化声音检测、方位角和仰角估计。在多种混响噪声条件下的大量实验表明,AuralNet优于近期方法。

原文摘要 · Abstract (English)

We propose AuralNet, a novel 3D multi-source binaural sound source localization approach that localizes overlapping sources in both azimuth and elevation without prior knowledge of the number of sources. AuralNet employs a gated coarse-tofine architecture, combining a coarse classification stage with a fine-grained regression stage, allowing for flexible spatial resolution through sector partitioning. The model incorporates a multi-head self-attention mechanism to capture spatial cues in binaural signals, enhancing robustness in noisy-reverberant environments. A masked multi-task loss function is designed to jointly optimize sound detection, azimuth, and elevation estimation. Extensive experiments in noisy-reverberant conditions demonstrate the superiority of AuralNet over recent methods

声源定位双耳感知注意力机制三维定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。