通过调整扬声器输出,降低语音识别设备周围的噪音,提升语音识别效果。
Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications
- 利用人耳听觉感知设计失真度量,动态调节扬声器输出以形成低音能区域。
- 实测与仿真均显示,在各类噪声场景下语音识别准确率显著提升。
- 可在牺牲部分听感质量的前提下,进一步降低设备周边声能,适合嵌入式语音系统。
本文提出一种鲁棒的扬声器波束成形算法,用于在扬声器产生主要噪声的场景(如音乐大声播放时)增强语音驱动应用(VDA)的语音识别性能。该算法通过调整扬声器播放信号,在实施语音识别的设备周围形成低声能区域。其基于人耳听觉感知设计失真度量,限制听者感知到的畸变。仿真与真实实验结果表明,所提方法在所有测试场景中均有效提升了语音识别性能。此外,该算法可在听者位置客观音质下降的代价下,进一步降低设备附近的声能密度。
原文摘要 · Abstract (English)
In this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing loudly. The loudspeaker beamformer modifies the loudspeaker playback signals to create a low-acoustic-energy region around the device that implements automatic speech recognition for a voice driven application (VDA). The algorithm utilises a distortion measure based on human auditory perception to limit the distortion perceived by human listeners. Simulations and real-world experiments show that the proposed loudspeaker beamformer improves the speech recognition performance in all tested scenarios. Moreover, the algorithm allows to further reduce the acoustic energy around the VDA device at the expense of reduced objective audio quality at the listener's location.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。