arXiv:2409.14346eess.AScs.SD2024-09被引 2

用可穿戴麦克风阵列提升动态环境下的语音方向估计精度

Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting

  • 引入可靠性加权机制,动态调整不同语音源的估计权重
  • 新质量度量使算法在混响噪声中仍能准确识别主方向
  • 适用于佩戴式设备在复杂移动场景中的语音定位

在多说话人房间环境中进行方向估计是众多应用的关键任务。尤其在说话人移动、混响和噪声共存的动态环境中,现有方法性能显著下降。本文基于可穿戴眼镜麦克风阵列采集的EasyCom数据集,研究改进局部空间域距离(LSDD)算法在嘈杂、动态、混响环境下的多说话人方向估计表现。原版LSDD算法在静态环境下表现优异,但在EasyCom动态数据中性能明显退化。通过全面的性能与系统分析,本文提出多项改进:引入可靠性加权策略,并设计新的聚类质量度量,有效识别更准确的方向估计结果,显著提升算法在复杂条件下的鲁棒性与准确性。

原文摘要 · Abstract (English)

Direction-of-arrival estimation of multiple speakers in a room is an important task for a wide range of applications. In particular, challenging environments with moving speakers, reverberation and noise, lead to significant performance degradation for current methods. With the aim of better understanding factors affecting performance and improving current methods, in this paper multi-speaker direction-of-arrival (DOA) estimation is investigated using a modified version of the local space domain distance (LSDD) algorithm in a noisy, dynamic and reverberant environment employing a wearable microphone array. This study utilizes the recently published EasyCom speech dataset, recorded using a wearable microphone array mounted on eyeglasses. While the original LSDD algorithm demonstrates strong performance in static environments, its efficacy significantly diminishes in the dynamic settings of the EasyCom dataset. Several enhancements to the LSDD algorithm are developed following a comprehensive performance and system analysis, which enable improved DOA estimation under these challenging conditions. These improvements include incorporating a weighted reliability approach and introducing a new quality measure that reliably identifies the more accurate DOA estimates, thereby enhancing both the robustness and accuracy of the algorithm in challenging environments.

方向估计可穿戴设备语音定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。