arXiv:2607.23027eess.AS2026-07

首次系统研究新加坡英语口音语音合成,微调后更像地道新加坡口音。

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

论文配图:Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
图 1 · 摘自论文原文
  • 用50位新加坡人语料微调现成语音模型,保留口音特征。
  • 微调后生成语音在口音相似度上显著提升,跨说话人也有效。
  • 适合想做方言/口音语音合成的研究者与开发者。

零样本文本转语音(ZS-TTS)对标准英语已接近人类水平,但对地域口音复制效果差。现有先进系统在输入短句式新加坡英语(Singlish)时,虽能保留说话人音色,却会弱化口音趋向通用英语。本文探究是否可通过微调现有ZS-TTS模型提升新加坡英语口音表现。我们在IMDA国家语音语料库的50位新加坡说话人数据上,对两款顶尖模型Chatterbox和CosyVoice 3进行微调。评估三种语音分布:真实录音、原始生成与微调后生成,均基于相同Singlish音频提示。评估涵盖自然度、可懂度、说话人相似性与口音相似性四个维度。区分了微调中包含的说话人(in-domain)与未见说话人(out-of-domain),测试口音迁移是否具有泛化能力。结果表明,微调显著提升了两类说话人在口音相似度上的表现,生成语音分布明显向真实新加坡英语靠拢,且该提升在未见说话人上依然成立。据我们所知,这是首个系统研究新加坡英语口音语音合成的工作。

原文摘要 · Abstract (English)

Zero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. Prompted with a short Singlish utterance, state-of-the-art systems reproduce a speaker's timbre while flattening the accent toward generic English. We investigate whether targeted fine-tuning off-the-shelf ZS-TTS can close the gap for Singapore English (Singlish). We fine-tune two cutting-edge ZS-TTS models, Chatterbox and CosyVoice 3, on 50 Singlish speakers from the IMDA National Speech Corpus. Three speech distributions are evaluated: real recordings against off-the-shelf and fine-tuned generation driven by the same Singlish audio prompts. The evaluation covers four dimensions: naturalness, intelligibility, speaker similarity, and accent similarity. We separate adaptation (in-domain speakers seen during fine-tuning) from consistency (held-out speakers) to test whether accent transfer generalises beyond the training data. Fine-tuning raises accent similarity on in-domain and out-of-domain speakers for both Chatterbox and CosyVoice 3. It moves the generated distribution measurably toward real Singlish, with the gain persisting on held-out speakers. To our knowledge, this is the first systematic study of Singlish-accented TTS.

语音合成口音建模微调新加坡英语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。