arXiv:2505.05056cs.CLcs.AI2025-05

首个带正字标注的潮州话野外语音数据集,助力低资源语言语音研究。

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations

  • 构建包含18.9小时潮州话语音的野外数据集,支持正式与口语表达。
  • 提供精确正字与拼音标注,适用于语音识别与语音合成任务。
  • 首次公开可用的潮州话正字标注数据集,适合方言语音研究者使用。

本文报道了潮州话野外语音语料库Teochew-Wild的构建。该语料库包含来自多位说话者的18.9小时潮州话语音数据,涵盖正式与非正式表达,并配有精确的正字与拼音标注。此外,我们还提供了配套的文本处理工具与资源,以推动该低资源语言在自动语音识别(ASR)与文本转语音(TTS)等任务中的研究与应用。据我们所知,这是首个公开可用且具有准确正字标注的潮州话语音数据集。我们在该语料库上进行了实验,结果验证了其在ASR与TTS任务中的有效性。

原文摘要 · Abstract (English)

This paper reports the construction of the Teochew-Wild, a speech corpus of the Teochew dialect. The corpus includes 18.9 hours of in-the-wild Teochew speech data from multiple speakers, covering both formal and colloquial expressions, with precise orthographic and pinyin annotations. Additionally, we provide supplementary text processing tools and resources to propel research and applications in speech tasks for this low-resource language, such as automatic speech recognition (ASR) and text-to-speech (TTS). To the best of our knowledge, this is the first publicly available Teochew dataset with accurate orthographic annotations. We conduct experiments on the corpus, and the results validate its effectiveness in ASR and TTS tasks.

语音识别方言数据集低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。