arXiv:2608.24659eess.AS2026-08

一个模型搞定未知说话人和麦克风数量的语音分离

REDnet: Recursive Encoder and Decoder for Speech Separation under Unknown Number of Speakers and Variable Number of Microphones

论文配图:REDnet: Recursive Encoder and Decoder for Speech Separation under Unknown Number of Speakers and Variable Number of Microphones
图 1 · 摘自论文原文
  • 递归编码解码结构,逐个分离说话人并融合空间信息
  • 在多个公开数据集上达到当前最优性能
  • 适合复杂多变麦克风阵列场景的语音分离任务

我们提出递归编码器与解码器(RED),构建一个单一深度神经网络(DNN)模型,用于分离包含未知数量说话人且麦克风数量可变、排列几何未知的多说话人混合语音,这一任务此前尚未被研究。RED 的解码器递归检测是否仍有活跃说话人,并逐次分离一人;其设计支持端到端训练以提升分离效果。编码器递归地对输入混合信号的每个麦克风通道进行编码,依次融合空间线索。结合两者,该 DNN 可训练实现对未知说话人数和可变麦克风数混合信号的分离,在多个公开数据集上取得当前最优性能。

原文摘要 · Abstract (English)

We propose $\textit{recursive encoder and decoder}$ (RED) for building a single deep neural network (DNN) model that can separate multi-speaker mixtures containing unknown numbers of speakers and variable numbers of microphones arranged in an unknown geometry, a task that has not been studied yet. The decoder of RED recursively detects whether there are active speakers left and separates one speaker at a time. It is designed to be trained in an end-to-end fashion to improve separation performance. The encoder of RED recursively encodes each microphone channel of the input mixture, sequentially incorporating spatial cues. Combining both, the DNN can be trained to separate mixtures not only with unknown numbers of speakers but also with variable numbers of microphones, achieving state-of-the-art performance on multiple public datasets.

语音分离多说话人自适应麦克风

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。