无需重训练,可直接修正语音语调和发音错误。
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
- 通过修改预训练模型内部表征实现后验控制。
- 有效调整语调并纠正发音错误,保持合成质量。
- 适合需要快速优化语音输出的场景。
近期文本转语音(TTS)技术显著提升了语音自然度,对精确语调控制和发音纠错的需求随之增加。现有语调调控方法通常依赖专用模块或额外训练,限制了其后验调整能力;传统发音纠错依赖音素转换字典,在低资源场景下实用性不足。本文提出反事实激活编辑(Counterfactual Activation Editing),一种与模型无关的方法,通过操纵预训练TTS模型内部表示,实现对语调和发音的后验控制。实验表明,该方法能有效调整语调特征并纠正发音错误,同时保持语音合成质量。这为无需重训练即可在推理阶段优化TTS输出提供了可能,弥合了预训练TTS模型与可编辑语音合成之间的差距。
原文摘要 · Abstract (English)
Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation often depend on specialized modules or additional training, limiting their capacity for post-hoc adjustments. Similarly, traditional mispronunciation correction relies on grapheme-to-phoneme dictionaries, making it less practical in low-resource settings. We introduce Counterfactual Activation Editing, a model-agnostic method that manipulates internal representations in a pre-trained TTS model to achieve post-hoc control of prosody and pronunciation. Experimental results show that our method effectively adjusts prosodic features and corrects mispronunciations while preserving synthesis quality. This opens the door to inference-time refinement of TTS outputs without retraining, bridging the gap between pre-trained TTS models and editable speech synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。