让神经网络学会在多人说话时聚焦目标说话人。
Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios
- 基于听众注视角度设计选择机制,指导模型训练时聚焦目标说话人。
- 相比传统方法,语音可懂度、音质和失真指标显著提升。
- 仅需音频输入推理,适合真实场景的语音增强应用。
我们提出一种针对端到端波束成形神经网络训练的说话人选择机制(SSM),基于最新发现:听者通常以一定欠角注视目标说话人。该机制使神经网络在多说话人场景下,根据听者与说话人的相对位置,学习聚焦于目标说话人。推理阶段仅需音频信息。通过声学仿真验证了该方法的可行性与性能。结果表明,与最小方差无失真滤波器及未使用SSM训练的同一神经网络相比,语音可懂度、音质和失真指标均有显著提升。该方法为解决鸡尾酒会问题迈出了重要一步。
原文摘要 · Abstract (English)
We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows the neural network model to learn toward which speaker to focus, during training, in a multi-speaker scenario, based on the position of listener and speakers. However, only audio information is necessary during inference. We perform acoustic simulations demonstrating the feasibility and performance when the SSM is employed in training. The results show significant increase in speech intelligibility, quality, and distortion metrics when compared to the minimum variance distortionless filter and the same neural network model trained without SSM. The success of the proposed method is a significant step forward toward the solution of the cocktail party problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。