针对孟加拉语语音识别与说话人分离,构建高质量数据流水线并实现高性能模型。
ShobdoSetu: A Data-Centric Framework for Bengali Long-Form Speech Recognition and Speaker Diarization
- 用抖音有声书和剧集构建高质量训练数据,结合大模型语言归一化与噪声增强。
- 语音识别在2.1万条数据上实现公榜WER 16.75、私榜15.55,说话人分离达公榜DER 0.1997。
- 适合资源稀缺语言的语音处理研究者,尤其关注数据工程与小样本适配。
孟加拉语使用者超2.3亿,但在自动语音识别(ASR)和说话人分离研究中严重不足。本文介绍我们在DL Sprint 4.0孟加拉语长语音识别(任务1)与说话人分离挑战(任务2)中的系统方案。任务1采用以数据为中心的流水线,从孟加拉语抖音有声书与戏剧中构建高质量训练语料,包含大模型辅助语言归一化、基于模糊匹配的片段边界验证及掩蔽区增强。在约21,000个数据点上微调tttugstugi/whisper-medium模型(束搜索大小为5),公榜词错误率(WER)达16.751,私榜为15.551。任务2在极端低资源设置(仅10个训练文件)下,通过针对性超参数优化微调pyannote.audio community-1分割模型,公榜说话人分离错误率(DER)为0.19974,私榜为0.26723。结果表明,精心的数据工程与领域自适应微调可在无大规模标注语料情况下实现竞争力表现。
原文摘要 · Abstract (English)
Bengali is spoken by over 230 million people yet remains severely under-served in automatic speech recognition (ASR) and speaker diarization research. In this paper, we present our system for the DL Sprint 4.0 Bengali Long-Form Speech Recognition (Task~1) and Bengali Speaker Diarization Challenge (Task~2). For Task~1, we propose a data-centric pipeline that constructs a high-quality training corpus from Bengali YouTube audiobooks and dramas \cite{tabib2026bengaliloop}, incorporating LLM-assisted language normalization, fuzzy-matching-based chunk boundary validation, and muffled-zone augmentation. Fine-tuning the \texttt{tugstugi/whisper-medium} model on approximately 21,000 data points with beam size 5, we achieve a Word Error Rate (WER) of 16.751 on the public leaderboard and 15.551 on the private test set. For Task~2, we fine-tune the pyannote.audio community-1 segmentation model with targeted hyperparameter optimization under an extreme low-resource setting (10 training files), achieving a Diarization Error Rate (DER) of 0.19974 on the public leaderboard, and .26723 on the private test set. Our results demonstrate that careful data engineering and domain-adaptive fine-tuning can yield competitive performance for Bengali speech processing even without large annotated corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。