STCON系统提升远场语音转写与说话人分离效果,关键在优化说话人归因与声源分离。
STCON System for the CHiME-8 Challenge
- 构建精细调校的说话人归因流水线,提升语音分割可靠性。
- 通过G-TSE模型降低错误率,实现更精准的声源分离。
- 适用于多麦克风环境下的远场语音识别任务。
本文介绍STCON系统在CHiME-8挑战赛任务1(DASR)中的应用,旨在利用多录音设备实现远场自动语音转写与说话人归因。研究重点在于精心训练与调优的说话人归因流水线及说话人数量估计,显著降低了说话人归因错误率(DER),并为语音分离与识别提供了更可靠的语音段。为提升声源分离性能,设计了引导目标说话人提取(G-TSE)模型,并与传统引导源分离(GSS)方法结合使用。为训练系统各模块,探索了多种数据增强与生成技术,有效提升了整体系统质量。
原文摘要 · Abstract (English)
This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trained and tuned diarization pipeline and speaker counting. This allowed to significantly reduce diarization error rate (DER) and obtain more reliable segments for speech separation and recognition. To improve source separation, we designed a Guided Target speaker Extraction (G-TSE) model and used it in conjunction with the traditional Guided Source Separation (GSS) method. To train various parts of our pipeline, we investigated several data augmentation and generation techniques, which helped us to improve the overall system quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。