arXiv:2411.06667eess.AScs.SD2024-11被引 8

融合说话人分离与聚类,提升复杂环境下语音识别准确率

DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions

  • 分步串联说话人辨识与语音分离模块,利用说话人时间边界优化分离效果
  • 在CHiME-8挑战赛中获单通道真实场景第一名,在LibriCSS上达新纪录
  • 引入窗口级解码和重聚类机制,有效缓解稀疏数据训练不稳问题

我们提出一种单通道深度级联融合说话人辨识与分离(DCF-DS)框架,用于后端自动语音识别(ASR)。该框架将神经说话人辨识(NSD)与语音分离(SS)模块顺序集成于联合训练框架中,使分离模块能有效利用辨识模块提供的说话人时间边界。为缓解稀疏数据收敛不稳(SDCI)问题,引入窗口级解码方案。还探索使用真实数据训练的NSD系统以获取更精准的说话人边界。此外,框架中可选加入多输入多输出语音增强模块(MIMO-SE),进一步提升性能。最后通过重聚类DCF-DS输出改进辨识结果,提升ASR精度。采用DCF-DS方法在CHiME-8 NOTSOFAR-1挑战赛的真实单通道赛道取得第一名,并在开放的LibriCSS数据集上达到当前最优单通道语音识别性能。

原文摘要 · Abstract (English)

We propose a single-channel Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework for back-end automatic speech recognition (ASR), combining neural speaker diarization (NSD) and speech separation (SS). First, we sequentially integrate the NSD and SS modules within a joint training framework, enabling the separation module to leverage speaker time boundaries from the diarization module effectively. Then, to complement DCF-DS training, we introduce a window-level decoding scheme that allows the DCF-DS framework to handle the sparse data convergence instability (SDCI) problem. We also explore using an NSD system trained on real datasets to provide more accurate speaker boundaries. Additionally, we incorporate an optional multi-input multi-output speech enhancement module (MIMO-SE) within the DCF-DS framework, which offers further performance gains. Finally, we enhance diarization results by re-clustering DCF-DS outputs, improving ASR accuracy. By incorporating the DCF-DS method, we achieved first place in the realistic single-channel track of the CHiME-8 NOTSOFAR-1 challenge. We also perform the evaluation on the open LibriCSS dataset, achieving a new state-of-the-art single-channel speech recognition performance.

语音识别说话人分离多说话人端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。