arXiv:2509.21382eess.AScs.SD2025-09被引 2

用说话人数量信息提升双耳助听器语音方向估计精度

Multi-Speaker DOA Estimation in Binaural Hearing Aids using Deep Learning and Speaker Count Fusion

  • 将说话人数量作为辅助特征融入深度学习模型
  • 真实数量信息使方向估计准确率最高提升14%
  • 适合研究助听设备语音增强与多源定位的学者

为在嘈杂多说话人环境中提取目标语音,双耳助听器中的到达方向(DOA)估计至关重要。现有方法中,基于麦克风信号谱相位差与幅值比的卷积循环神经网络(CRNN)是常用方案。本文探索在多说话人场景下引入源数量信息:首先采用联合多源DOA估计与说话人计数的双任务训练,随后在独立的DOA估计系统中,通过早期、中期和晚期融合策略将说话人数量(0、1或2+个)作为辅助特征嵌入CRNN结构。基于真实双耳录音的实验表明,双任务训练虽提升计数性能,但未改善DOA估计;而使用真实(即‘理想’)说话人数量作为辅助特征时,独立系统性能显著提升,晚期融合可使平均F1分数相比基线CRNN最高提高14%。这表明利用源数量估计可显著增强双耳助听器中稳健的DOA估计能力。

原文摘要 · Abstract (English)

For extracting a target speaker voice, direction-of-arrival (DOA) estimation is crucial for binaural hearing aids operating in noisy, multi-speaker environments. Among the solutions developed for this task, a deep learning convolutional recurrent neural network (CRNN) model leveraging spectral phase differences and magnitude ratios between microphone signals is a popular option. In this paper, we explore adding source-count information for multi-sources DOA estimation. The use of dual-task training with joint multi-sources DOA estimation and source counting is first considered. We then consider using the source count as an auxiliary feature in a standalone DOA estimation system, where the number of active sources (0, 1, or 2+) is integrated into the CRNN architecture through early, mid, and late fusion strategies. Experiments using real binaural recordings are performed. Results show that the dual-task training does not improve DOA estimation performance, although it benefits source-count prediction. However, a ground-truth (oracle) source count used as an auxiliary feature significantly enhances standalone DOA estimation performance, with late fusion yielding up to 14% higher average F1-scores over the baseline CRNN. This highlights the potential of using source-count estimation for robust DOA estimation in binaural hearing aids.

语音增强方向估计助听器深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。