让文字更懂声音:SPARCLE提升低资源语音合成质量
SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

- 用对比学习对齐字符与声学特征,结合说话人信息生成声学感知的字符表示
- 在极端低资源场景下,词错误率降低50%,显著优于传统字符模型
- 适合低资源语音合成、个性化语音生成等场景,尤其关注说话人差异
语音合成近期从音素表征转向直接的字形建模。尽管音素能处理文本到声音的一对多映射,但依赖的音形转换(G2P)系统无法捕捉说话人特有的声学差异。已有研究表明,字形模型在大规模数据上表现更优,但在低资源环境下表现不佳。本文提出SPARCLE,一种考虑说话人的字形表征模型,通过对比学习将字形与其对应的Wav2Vec2声学表示对齐,并以说话人身份为条件。该模型可替代传统G2P系统用于下游文语转换任务。实验表明,SPARCLE在极端低资源设置下将词错误率降低50%,显著提升生成质量。
原文摘要 · Abstract (English)
Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P) systems that fail to capture speaker-specific acoustic variation. Prior work demonstrates that grapheme-based models outperform phoneme-based systems at scale, but not in low-resource settings. In this paper, we propose SPARCLE, a speaker-aware grapheme representation model that enriches characters with their precise acoustic realizations. SPARCLE is trained with a contrastive objective to align graphemes with corresponding Wav2Vec2 acoustic representations while conditioned on speaker identity. The resulting model serves as a replacement to G2P systems for downstream text-to-speech (TTS) tasks. We demonstrate that SPARCLE improves generation quality, reducing word error rates by half in extreme low-resource settings compared to standard grapheme-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。