用生成流网络减少大模型语音合成中的幻觉,无需额外训练资源。
Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
- 将语音生成转为轨迹流优化,通过不确定性分析指导修正。
- 测试集字符错误率降低超50%,模型不确定性下降最多58%。
- 适合追求低延迟、高准确的语音合成系统部署者使用。
基于语言模型(LM)的文本转语音(TTS)系统常生成偏离输入文本的幻觉语音。现有缓解策略或需大量训练资源,或引入显著推理延迟。本文提出基于生成流网络的分布对齐框架GOAT,一种无需大量资源与推理开销的后训练方法。我们首先进行不确定性分析,发现幻觉与模型不确定性呈强正相关。据此,将TTS生成重构为轨迹流优化问题,引入改进的子轨迹平衡目标及强化的内部奖励作为目标分布,并结合奖励温度衰减与学习率优化以提升稳定性与性能平衡。大量实验表明,GOAT在挑战性测试案例上使字符错误率降低超过50%,不确定性降低达58%,展现出优异的泛化能力与有效性。
原文摘要 · Abstract (English)
Language Model (LM)-based Text-to-Speech (TTS) systems often generate hallucinated speech that deviates from input text. Existing mitigation strategies either demand excessive training resources or introduce significant inference latency. In this paper, we propose GFlOwNet-guided distribution AlignmenT (GOAT) for LM-based TTS, a post-training framework that mitigates hallucinations without relying on massive resources or inference cost. Specifically, we first conduct an uncertainty analysis, revealing a strong positive correlation between hallucination and model uncertainty. Based on this, we reformulate TTS generation as a trajectory flow optimization problem and introduce an enhanced Subtrajectory Balance objective together with a sharpened internal reward as target distribution. We further integrate reward temperature decay and learning rate optimization for stability and performance balance. Extensive experiments show that GOAT reduce over 50% character error rates on challenging test cases and lowering uncertainty by up to 58%, demonstrating its strong generalization ability and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。