arXiv:2409.16803eess.AScs.SD2024-09被引 5

利用多通道空间信息提升语音分离准确率,让说话人辨识更精准。

Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings

  • 分三阶段用多通道空间信息优化单通道语音分离流程
  • 在CHiME-8挑战中获第一名,显著降低说话人错误率
  • 适合需要高精度语音分离的远场会议场景

尽管端到端说话人辨识系统近年进展显著,但在真实场景中,模块化系统因更强的适应性和鲁棒性仍表现更优。传统模块化方法很少利用多通道语音中的空间线索。本文提出一种三阶段模块化系统,通过多通道语音的空间信息提升单通道神经说话人辨识(NSD)的初始化准确性:(1) 在多通道语音上进行重叠检测与连续语音分离(CSS),获取更清晰的单说话人片段用于聚类,并执行首次NSD解码;(2) 利用首次解码结果初始化复杂角中心高斯混合模型(cACGMM),估计说话人级掩码,结合重叠相加与掩码转语音活动检测(Mask-to-VAD),实现更低说话人错误率(SpkErr)的初始化,随后进行第二次NSD解码;(3) 使用第二次解码结果指导源分离(GSS),识别并过滤包含少于一个词的短片段,获得更干净语音,再进行重聚类和最终的NSD解码。我们在CHiME-8 NOTSOFAR-1(远场录音下的自然办公室对话者)挑战赛中展示了逐步探索的评估结果,验证了系统的有效性及其对识别性能的提升。最终系统在挑战赛中排名第一。

原文摘要 · Abstract (English)

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically, modular speaker diarization methods have seldom discussed how to leverage spatial cues from multi-channel speech. This paper proposes a three-stage modular system to enhance single-channel neural speaker diarization systems and recognition performance by utilizing spatial cues from multi-channel speech to provide more accurate initialization for each stage of neural speaker diarization (NSD) decoding: (1) Overlap detection and continuous speech separation (CSS) on multi-channel speech are used to obtain cleaner single speaker speech segments for clustering, followed by the first NSD decoding pass. (2) The results from the first pass initialize a complex Angular Central Gaussian Mixture Model (cACGMM) to estimate speaker-wise masks on multi-channel speech, and through Overlap-add and Mask-to-VAD, achieve initialization with lower speaker error (SpkErr), followed by the second NSD decoding pass. (3) The second decoding results are used for guided source separation (GSS), recognizing and filtering short segments containing less one word to obtain cleaner speech segments, followed by re-clustering and the final NSD decoding pass. We presented the progressively explored evaluation results from the CHiME-8 NOTSOFAR-1 (Natural Office Talkers in Settings Of Far-field Audio Recordings) challenge, demonstrating the effectiveness of our system and its contribution to improving recognition performance. Our final system achieved the first place in the challenge.

说话人辨识多通道语音空间线索语音分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。