arXiv:2605.17846eess.AS2026-05

构建156小时高保真乌尔都语语音库,含12维副语言标注,推动低资源语言AI发展。

UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations

  • 用大模型驱动管道,覆盖新闻、戏剧等12类内容,解决右向左书写与混用问题
  • 主数据集156小时,男女比例60:40,经人工评估均分4.6,可信度达97.6%
  • 公开9小时标准标注集,适合做乌尔都语语音研究与多模态任务的开发者

尽管有2.3亿使用者,乌尔都语在语音技术方面仍严重缺乏资源。我们提出UrduSpeech:一个包含156小时音频的大型高质量乌尔都语语料库,附带12维副语言元数据,涵盖标准语(US-Std)、口语(US-CS)及英式巴基斯坦语(US-EngPk)。为应对右向左书写限制和频繁代码切换,开发了基于大模型的语料采集管道,覆盖新闻、戏剧及罕见文学形式如Bait-Bazi等12个类别。同时发布9小时经过母语者人工校正的US-Benchmark集作为基准。对主156小时语料的人工质量评估显示平均意见分(MOS)为4.6(标准差=0.7),评者间一致性经Cohen's Kappa确认为0.68,验证了数据清洗管道的97.6%置信度。语料库共含71,792条语音片段,性别比例保持60:40。本工作显著推进全球AI中的语言包容性。语料与代码已开源,提供演示页面。

原文摘要 · Abstract (English)

Despite 230 million speakers, Urdu remains critically under-resourced in speech technology. We introduce UrduSpeech: a large high-fidelity Urdu corpus comprising 156 hours of audio with 12-dimension paralinguistic metadata, encompassing US-Std, US-CS, US-EngPk. To address Right-to-Left script constraints and frequent code-switching, we developed UrduSpeech, a LLM-driven pipeline to curate data across 12 diverse categories, including news, drama, and rare literary forms like Bait-Bazi. We also release a 9-hour US-Benchmark set, manually corrected by native annotators to serve as a standard. Human quality assessment of the primary 156-hour corpus yielded a Mean Opinion Score (MOS) of 4.6 (std = 0.7) with inter-rater reliability confirmed by a 0.68 Cohen's Kappa, validating our curation pipeline's 97.6% confidence score. The corpus maintains a 60-40 gender balance across 71,792 utterances. Our work represents a significant leap toward linguistic inclusivity in global AI. The corpus and code are open-sourced, and a demo page is available.

语音数据集多模态乌尔都语低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。