端到端模型直接输出清晰立体语音,提升远程会议听感。
End-to-end multi-channel speaker extraction and binaural speech synthesis
- 统一提取、降噪与立体声渲染的端到端深度学习框架
- 在多通道混响噪声下显著提升语音质量和空间保真度
- 适合需要高质量语音沉浸体验的远程会议系统
语音清晰度和空间音频沉浸感是提升远程会议体验的两大关键因素。现有方法往往受限:单麦克风缺乏空间信息,而麦克风阵列方法性能高度依赖方向估计精度。为此,我们提出一种端到端深度学习框架,可直接将多通道嘈杂混响信号映射为清晰且具有空间感的双耳语音。该框架将声源分离、降噪与双耳渲染统一于一个网络中。提出一种新型幅度加权的双耳级差损失函数,旨在提升空间还原精度。大量实验表明,该方法在语音质量与空间保真度上均优于现有基线。
原文摘要 · Abstract (English)
Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or because their performance is highly dependent on the accuracy of direction-of-arrival estimation when using microphone array. To overcome this issue, we introduce an end-to-end deep learning framework that has the capacity of mapping multi-channel noisy and reverberant signals to clean and spatialized binaural speech directly. This framework unifies source extraction, noise suppression, and binaural rendering into one network. In this framework, a novel magnitude-weighted interaural level difference loss function is proposed that aims to improve the accuracy of spatial rendering. Extensive evaluations show that our method outperforms established baselines in terms of both speech quality and spatial fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。