让语音合成自由切换语调、情绪,不依赖特定参考音频。
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
- 分离文本与参考音色的引导,减少风格依赖。
- 支持连续调节语调、能量和多种情绪,保持发音清晰。
- 适合需要灵活控制语音风格的研究者和开发者。
零样本文语转换模型可从短参考音频中克隆说话人音色,但也会继承参考音频中的表达风格,导致在仅有限或不匹配参考音频时难以合成所需风格。现有可控语音合成方法多依赖绝对风格目标或离散文本提示,无法实现连续且参考相对的风格控制。本文提出 ReStyle-TTS,通过解耦分类器无关引导(DCFG),独立控制文本与参考引导,降低对参考风格的依赖;结合风格专用 LoRA 与正交 LoRA 融合,实现连续、解耦的多属性控制,并引入音色一致性优化模块缓解弱化参考引导带来的音色漂移。实验表明,ReStyle-TTS 可实现用户友好的连续、相对风格控制,在语调、能量和多种情绪上表现良好,同时保持语音可懂度与说话人音色一致,且在挑战性的参考-目标风格不匹配场景下仍具鲁棒性。
原文摘要 · Abstract (English)
Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result, synthesizing speech with a desired style often requires carefully selecting reference audio, which is impractical when only limited or mismatched references are available. While recent controllable TTS methods attempt to address this issue, they typically rely on absolute style targets and discrete textual prompts, and therefore do not support continuous and reference-relative style control. We propose ReStyle-TTS, a framework that enables continuous and reference-relative style control in zero-shot TTS. Our key insight is that effective style control requires first reducing the model's implicit dependence on reference style before introducing explicit control mechanisms. To this end, we introduce Decoupled Classifier-Free Guidance (DCFG), which independently controls text and reference guidance, reducing reliance on reference style while preserving text fidelity. On top of this, we apply style-specific LoRAs together with Orthogonal LoRA Fusion to enable continuous and disentangled multi-attribute control, and introduce a Timbre Consistency Optimization module to mitigate timbre drift caused by weakened reference guidance. Experiments show that ReStyle-TTS enables user-friendly, continuous, and relative control over pitch, energy, and multiple emotions while maintaining intelligibility and speaker timbre, and performs robustly in challenging mismatched reference-target style scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。