arXiv:2607.02633cs.LGcs.CL2026-07

用短音频样本精准控制单词发音,解决零样本语音合成中的生僻词发音难题。

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

论文配图:GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
图 1 · 摘自论文原文
  • 通过编码参考音频并绑定到目标词位置,实现逐词发音控制。
  • 在五种语言上降低22%-39%的音素错误率,对难词发音更接近参考录音。
  • 适合需要高精度发音控制的应用,如方言、专有名词或语音克隆场景。

我们提出GRAFT,一种用于文本到语音神经编解码语言建模的逐词发音控制机制。现有系统虽具备高可懂性和自然度,但受制于文本歧义,常误读罕见专有名词、外来词和技术术语。即使采用音素条件建模,也缺乏对单个词汇发音的直接声学调控能力。GRAFT通过将目标词的简短语音样本编码并绑定至提示中的对应位置,实现精确发音控制。训练数据构建时使用语音转换技术,使提示说话人与目标说话人分离,因此提示可来自任意声音,而输出仍保持目标音色。在盲听英语测试中,人类评估者一致认为GRAFT表现最佳,其对困难词汇的发音最接近参考录音。在五语言客观基准测试中,GRAFT相比仅文本输入的基线模型将目标词音素错误率降低22%-39%,优于其他开源零样本系统(包括音素和文本条件),同时保持了说话人相似性与自然度。

原文摘要 · Abstract (English)

We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but inherit the ambiguity of text and mispronounce rare proper nouns, loanwords and technical terms. Even phoneme-conditioned models offer no direct acoustic handle for per-word pronunciation. GRAFT controls the pronunciation of a chosen word from a short spoken sample of it, encoded with the model's own speech tokenizer and bound to the word's position in the prompt. Voice conversion during training-data construction disentangles the hint speaker from the target speaker, so the hint may come from any voice while the output stays in the target voice. In a blind English listening study, human raters rank GRAFT first by a clear margin, judging its rendering of the difficult word closest to a reference recording of that word. On a five-language objective benchmark, GRAFT reduces target-word phoneme error rate by 22-39% over the identical text-only backbone and outperforms competitive open-source zero-shot systems, both phoneme- and text-conditioned, on target-word pronunciation, while preserving speaker similarity and naturalness.

语音合成发音控制零样本语音克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。