arXiv:2509.18004cs.CLcs.SD2025-09被引 15

构建1万小时四川话语音语料库,助力方言语音技术研究

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

  • 基于自研管道框架构建大规模四川话语音数据集
  • 训练模型性能达开源系统顶尖水平,接近商用服务
  • 适合方言语音、AI公平性研究者使用

方言大规模开源数据匮乏严重制约语音技术发展,尤其在汉语四川话这类广泛使用的方言上。为解决这一关键问题,我们提出 WenetSpeech-Chuan,一个10,000小时、标注丰富的四川话语音语料库,基于自研的 Chuan-Pipeline 框架构建。为支持严谨评估,我们还发布了高质量的语音识别(ASR)与语音合成(TTS)基准测试集 WenetSpeech-Chuan-Eval,包含人工校验的转写文本。实验表明,基于该语料库训练的模型在开源系统中达到领先性能,且表现可媲美商业服务。作为目前最大的开源四川话语音语料库,WenetSpeech-Chuan 不仅降低了方言语音研究门槛,也对推动语音技术公平性、减少偏见具有重要意义。语料库、基准、模型及凭证均已公开发布于项目主页。

原文摘要 · Abstract (English)

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constructed using our novel Chuan-Pipeline, a complete data processing framework for dialectal speech. To facilitate rigorous evaluation and demonstrate the corpus's effectiveness, we also release high-quality ASR and TTS benchmarks, WenetSpeech-Chuan-Eval, with manually verified transcriptions. Experiments show that models trained on WenetSpeech-Chuan achieve state-of-the-art performance among open-source systems and demonstrate results comparable to commercial services. As the largest open-source corpus for Sichuanese dialects, WenetSpeech-Chuan not only lowers the barrier to research in dialectal speech processing but also plays a crucial role in promoting AI equity and mitigating bias in speech technologies. The corpus, benchmarks, models, and receipts are publicly available on our project page.

语音识别方言处理数据集四川话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。