提出一种无需训练的解码方法,有效减少语音合成中的幻觉问题。
Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech
- 通过对比有无文本条件下的预测,强化文本对齐支持
- 在4个模型上降低词错率最高达55.6%,24/25多语言场景表现提升
- 适合关注语音合成稳定性与自然度的研究者与开发者
基于语言模型的文本到语音(LM-based TTS)仍易出现偏离目标文本的语音幻觉。现有缓解方法主要依赖结构修改或额外训练,而解码阶段的控制仍被忽视。本文提出一种条件信息视角,区分源自文本的对齐信息与由声学上下文和学习到的语音规律提供的经验信息。假设一类幻觉发生在对齐支持不足的关键过渡点。通过同一语音语言模型在有无文本条件下的预测,提出无需训练的体验校准对比解码(ECCD),在增强对齐支持的同时保留有用的经验信息。ECCD保持原始专家分布,仅实施正向对齐增强,并使用集合级经验兼容性校准强度。在四个模型上,ECCD在所有SeedTTS-Eval设置中将词错率(WER)/字符错率(CER)降低最多达55.6%,在24/25个多语言CV3-Eval设置中表现提升。听感测试显示客观质量评分(CMOS)提升+0.644,同时保持强说话人相似性。进一步分析表明,对齐影响和决策增益在语言单元内差异显著,首次错误边界处低于匹配正确边界处。整体实验与分析表明,条件信息控制是缓解语音幻觉的有前景的解码阶段方向。
原文摘要 · Abstract (English)
Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text. Existing mitigation mainly relies on architectural changes or additional training, while decoding-time control remains underexplored. We present a conditional information view that distinguishes text-derived alignment information from experience information supplied by acoustic context and learned speech regularities. We hypothesize that an important class of hallucinations begins when alignment support is insufficiently reflected in the selected token at a vulnerable transition. Using predictions from the same speech LM with and without text conditions, we propose Experience-Calibrated Contrastive Decoding (ECCD), a training-free method that strengthens alignment support while preserving useful experience information. ECCD preserves the original expert distribution, applies only positive alignment enhancement, and calibrates its strength using set-level experience compatibility. Across four models, ECCD reduces WER/CER by up to 55.6% in all SeedTTS-Eval settings and 24 of 25 multilingual CV3-Eval settings. A listening test yields a CMOS gain of $+0.644$ while retaining strong speaker similarity. Further analysis shows that alignment influence and decision-level gain vary within linguistic units and are lower at first-error boundaries than at matched correct boundaries. Overall, these extensive experiments and analyses identify conditional information control as a promising decoding-time direction for mitigating speech hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。