提出轻量级端到端双耳语音定位模型,提升嘈杂环境中的声源定位精度。
Binaural Localization Model for Speech in Noise
- 采用轻量卷积循环网络处理双耳噪声混响信号
- 在真实听觉阈值下实现前向方位角精确定位
- 可评估双耳语音增强方法的线索保留效果
双耳声源定位对人类的空间感知、沟通与安全至关重要。本文提出一种面向噪声环境的端到端双耳定位模型,采用轻量级卷积循环网络,对前方方位角平面内的噪声混响双耳信号进行定位。模型引入内耳加性噪声以模拟典型听者的频率依赖性听觉阈值。通过对比波束成形响应功率算法,评估了该模型的定位性能,并研究其作为双耳语音增强方法中双耳线索保留度量的适用性。还进行了听觉实验,比较该模型与人类在嘈杂条件下对语音定位的表现。
原文摘要 · Abstract (English)
Binaural acoustic source localization is important to human listeners for spatial awareness, communication and safety. In this paper, an end-to-end binaural localization model for speech in noise is presented. A lightweight convolutional recurrent network that localizes sound in the frontal azimuthal plane for noisy reverberant binaural signals is introduced. The model incorporates additive internal ear noise to represent the frequency-dependent hearing threshold of a typical listener. The localization performance of the model is compared with the steered response power algorithm, and the use of the model as a measure of interaural cue preservation for binaural speech enhancement methods is studied. A listening test was performed to compare the performance of the model with human localization of speech in noisy conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。