用目标说话人短语音引导,实现更准的多麦克风语音提取
MC-LExt: Multi-Channel Target Speaker Extraction with Onset-Prompted Speaker Conditioning Mechanism
- 在多通道混合信号前加目标说话人录音作为引导信号
- 在WHAMR!和MC-Libri2Mix上提升语音提取效果
- 无需方向或嵌入,适合嘈杂混响环境
多通道目标说话人提取(MC-TSE)旨在从多个麦克风捕获的多人语音信号中分离出特定说话人的声音。现有方法常依赖方向到达角(DOA)或说话人嵌入等辅助信息,但基于DOA的方法依赖显式方向估计且对麦克风阵列结构敏感,而基于嵌入的方法隐式建模说话人身份,在噪声混响条件下性能下降。为此,我们提出多通道听音提取(MC-LExt),一种简单高效的新框架。核心思想是在每个通道的多通道混合信号前添加一段目标说话人的短录语音,形成起始提示条件信号,引导语音提取。该设计使深度神经网络能端到端联合学习空间与说话人身份特征。在包含WHAMR!和MC-Libri2Mix的噪声混响基准测试中,验证了该方法的有效性。
原文摘要 · Abstract (English)
Multi-channel target speaker extraction (MC-TSE) aims to extract a target speaker's voice from multi-speaker signals captured by multiple microphones. Existing methods often rely on auxiliary clues such as direction-of-arrival (DOA) or speaker embeddings. However, DOA-based approaches depend on explicit direction estimation and are sensitive to microphone array geometry, while methods based on speaker embeddings model speaker identity in an implicit manner and may degrade in noisy-reverberant conditions. To address these limitations, we propose multi-channel listen to extract (MC-LExt), a simple but highly-effective framework for MC-TSE. Our key idea is to prepend a short enrollment utterance of the target speaker to each channel of the multi-channel mixture, providing an onset-prompted conditioning signal that can guide TSE. This design allows the deep neural network (DNN) to learn spatial and speaker identity cues jointly in a fully end-to-end manner. Experiments on noisy-reverberant benchmarks, including WHAMR! and MC-Libri2Mix, demonstrate the effectiveness of MC-TSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。