针对长时孟加拉语语音,提升识别与说话人分离的鲁棒性。
A Holistic Framework for Robust Bangla ASR and Speaker Diarization with Optimized VAD and CTC Alignment
- 优化语音活动检测与CTC强制词对齐,保持长时间语音时序准确
- 在超过3060秒的长音频上实现稳定识别与说话人分离
- 适合需要处理复杂多说话人场景的孟加拉语应用开发者
尽管孟加拉语是全球使用最广泛的语言之一,但在自然语言处理领域仍属低资源语言。主流的孟加拉语自动语音识别(ASR)与说话人分离系统在处理超过3060秒的长音频时表现不佳。本文针对DL Sprint 4.0竞赛,提出一个专为长时孟加拉语内容设计的鲁棒框架,基于现有模型并引入新型优化流水线。方法包括语音活动检测(VAD)优化和基于强制词对齐的连接时序分类(CTC)分段,以保障长时间语音中的时序精度与转录完整性。同时采用多种微调技术,并通过增强与降噪预处理数据。该工作有效弥合了复杂多说话人环境下的性能差距,为真实场景中的长时孟加拉语语音应用提供了可扩展解决方案。
原文摘要 · Abstract (English)
Despite being one of the most widely spoken languages globally, Bangla remains a low-resource language in the field of Natural Language Processing (NLP). Mainstream Automatic Speech Recognition (ASR) and Speaker Diarization systems for Bangla struggles when processing longform audio exceeding 3060 seconds. This paper presents a robust framework specifically engineered for extended Bangla content by leveraging preexisting models enhanced with novel optimization pipelines for the DL Sprint 4.0 contest. Our approach utilizes Voice Activity Detection (VAD) optimization and Connectionist Temporal Classification (CTC) segmentation via forced word alignment to maintain temporal accuracy and transcription integrity over long durations. Additionally, we employed several finetuning techniques and preprocessed the data using augmentation techniques and noise removal. By bridging the performance gap in complex, multi-speaker environments, this work provides a scalable solution for real-world, longform Bangla speech applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。