arXiv:2604.11803cs.CL2026-04中稿 · DialRes-LREC26

构建了六小时萨尔布吕肯德语方言语音语料库,助力方言语音合成研究。

Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus

论文配图:Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
图 1 · 摘自论文原文
  • 基于书籍与本地材料收集文本,九位说话人录制语音并校准
  • 提供对齐的文本与音频数据,支持零样本和少样本语音合成
  • 解决方言拼写与发音差异问题,为低资源方言建模提供基础

近年来自然语言处理与语音技术取得了显著进展,但主要聚焦于标准语言形式。尽管方言具有重要文化价值且广泛使用,却在语言资源与计算模型中严重缺位,导致性能差距。为此,我们推出萨尔布吕肯方言语音语料库 Saar-Voice,包含六小时语音数据。通过数字化书籍与本地资料收集文本,选取其中子集由九位说话人录制,并对文本与语音成分进行分析以评估数据特性与质量。论文讨论了拼写与说话人变异带来的方法论挑战,探索了字符到发音(G2P)转换。最终提供的语料库包含对齐的文本与音频表示,为未来方言感知的文本转语音(TTS)研究奠定基础,尤其适用于低资源场景下的零样本与少样本模型适配。

原文摘要 · Abstract (English)

Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread use, are underrepresented in linguistic resources and computational models, resulting in performance disparities. To address this gap, we introduce Saar-Voice, a six-hour speech corpus for the Saarbrücken dialect of German. The dataset was created by first collecting text through digitized books and locally sourced materials. A subset of this text was recorded by nine speakers, and we conducted analyses on both the textual and speech components to assess the dataset's characteristics and quality. We discuss methodological challenges related to orthographic and speaker variation, and explore grapheme-to-phoneme (G2P) conversion. The resulting corpus provides aligned textual and audio representations. This serves as a foundation for future research on dialect-aware text-to-speech (TTS), particularly in low-resource scenarios, including zero-shot and few-shot model adaptation.

语音合成方言识别低资源语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。