arXiv:2602.14291cs.SDeess.AS2026-02被引 8

构建首个孟加拉语长语音识别与说话人分离基准,推动本地化语音技术发展。

Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization

  • 基于YouTube内容构建158.6小时孟加拉语语音数据集,经人工校验确保质量
  • 建立34.07%词错误率(WER)的语音识别基线和40.08%段错误率(DER)的说话人分离基线
  • 提供标准化评估协议与标注规范,适合研究者开展可复现的孟加拉语语音研究

孟加拉语在长语音技术方面仍属资源匮乏,尽管其使用广泛。本文提出Bengali-Loop,两个社区基准:(1) 来自11个YouTube频道的191段录音(总计158.6小时,79.2万词)的长语音自动语音识别(ASR)语料库,通过可复现的字幕提取流程并结合人工审核验证;(2) 24段录音(共22小时,5,744个标注说话人片段)的说话人分离语料库,采用全手动标注的说话人片段标签,格式为CSV。两个数据集均针对真实场景下的多说话人、长时长内容(如孟加拉语戏剧/natok)。我们建立了基线模型(Tugstugi: 34.07% WER;pyannote.audio: 40.08% DER),并提供标准化评估协议(WER/CER、DER)、标注规则和数据格式,以支持可复现的基准测试及未来孟加拉语长语音识别与分离模型的开发。

原文摘要 · Abstract (English)

Bengali (Bangla) remains under-resourced in long-form speech technology despite its wide use. We present Bengali-Loop, two community benchmarks to address this gap: (1) a long-form ASR corpus of 191 recordings (158.6 hours, 792k words) from 11 YouTube channels, collected via a reproducible subtitle-extraction pipeline and human-in-the-loop transcript verification; and (2) a speaker diarization corpus of 24 recordings (22 hours, 5,744 annotated segments) with fully manual speaker-turn labels in CSV format. Both benchmarks target realistic multi-speaker, long-duration content (e.g., Bangla drama/natok). We establish baselines (Tugstugi: 34.07% WER; pyannote.audio: 40.08% DER) and provide standardized evaluation protocols (WER/CER, DER), annotation rules, and data formats to support reproducible benchmarking and future model development for Bangla long-form ASR and diarization.

语音识别说话人分离孟加拉语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。