arXiv:2602.23070cs.SDcs.AI2026-02

针对孟加拉语长音频,用增强与对齐提升语音识别与说话人分离效果。

Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment

  • 用合成噪声和混响增强数据,结合精确标注进行微调。
  • 在882小时多说话人数据上实现0.019实时因子的高效系统。
  • 适合低资源语言长语音处理研究者参考。

尽管孟加拉语自动语音识别(ASR)已取得显著进展,但处理长时音频和鲁棒说话人分离仍是关键研究空白。为解决该语言联合ASR与分离资源严重匮乏的问题,我们推出了882小时多说话人孟加拉语数据集Lipi-Ghor-882。本文详述了参加DL Sprint 4.0竞赛的方案,系统评估了多种长时孟加拉语语音模型架构。对于ASR,发现原始数据缩放无效;相反,利用精确对齐标注并结合合成声学退化(噪声与混响)的针对性微调成为最有效方法。而对于说话人分离,全球开源先进模型(如Diarizen)在该复杂数据集上表现意外不佳。大规模模型重训练改进甚微;相反,对基线模型输出进行策略性启发式后处理成为提升准确率的主要驱动力。最终,本工作构建出一个高度优化的双流管道,实现实时因子约0.019,为低资源、长时语音处理建立了可实践的实证基准。

原文摘要 · Abstract (English)

Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor-882, a comprehensive 882-hour multi-speaker Bengali dataset. In this paper, detailing our submission to the DL Sprint 4.0 competition, we systematically evaluate various architectures and approaches for long-form Bengali speech. For ASR, we demonstrate that raw data scaling is ineffective; instead, targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach. Conversely, for speaker diarization, we observed that global open-source state-of-the-art models (such as Diarizen) performed surprisingly poorly on this complex dataset. Extensive model retraining yielded negligible improvements; instead, strategic, heuristic post-processing of baseline model outputs proved to be the primary driver for increasing accuracy. Ultimately, this work outlines a highly optimized dual pipeline achieving a $\sim$0.019 Real-Time Factor (RTF), establishing a practical, empirically backed benchmark for low-resource, long-form speech processing.

语音识别说话人分离低资源语言数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。