arXiv:2505.22013cs.SDeess.AS2025-05中稿 · Interspeech 2025被引 2

融合分段与聚类,提升重叠语音场景下的说话人分离与识别准确率。

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge

  • 混合使用端到端分段与传统聚类,自适应处理不同重叠程度的语音。
  • 在低信噪比下通过感知ASR的观测增强,提升识别性能。
  • 端到端级联架构,实现在真实会议场景中的领先表现。

本文针对MISP 2025挑战赛构建了系统方案。说话人分离方面,提出结合WavLM端到端分段与多模块聚类的混合方法,根据重叠程度自适应选择模型。语音识别方面,设计了考虑ASR需求的观测增强策略,以弥补低信噪比下引导源分离(GSS)的性能不足。最终采用级联架构集成两系统以应对第三赛道。系统在第二赛道取得9.48%的字符错误率(CER),在第三赛道实现11.56%的拼接最小排列字符错误率(cpCER),双赛道均排名第一,验证了所提方法在真实会议场景中的有效性。

原文摘要 · Abstract (English)

This paper presents the system developed to address the MISP 2025 Challenge. For the diarization system, we proposed a hybrid approach combining a WavLM end-to-end segmentation method with a traditional multi-module clustering technique to adaptively select the appropriate model for handling varying degrees of overlapping speech. For the automatic speech recognition (ASR) system, we proposed an ASR-aware observation addition method that compensates for the performance limitations of Guided Source Separation (GSS) under low signal-to-noise ratio conditions. Finally, we integrated the speaker diarization and ASR systems in a cascaded architecture to address Track 3. Our system achieved character error rates (CER) of 9.48% on Track 2 and concatenated minimum permutation character error rate (cpCER) of 11.56% on Track 3, ultimately securing first place in both tracks and thereby demonstrating the effectiveness of the proposed methods in real-world meeting scenarios.

说话人分离语音识别重叠语音端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。