用少量标注数据实现角色语音的快速定制,让语音更自然有情感。
Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning
- 用文本风格标记+高质量音频提示,实现单次调整即适配新声音
- 通过上下文学习与在线强化学习,提升语音自然度与情感表达
- 适合需要快速生成个性化语音的场景,如游戏或虚拟助手
对话式人工智能进展显著,但生成富有表现力且可控的文本转语音(TTS)仍具挑战性,尤其是精细控制语音风格和情绪,通常需大量标注数据。为突破数据瓶颈,我们提出一种可扩展、数据高效的级联框架,将文本风格标记与人工精选的高质量音频提示相结合,实现对细粒度语调风格和角色声音的单次适应。在TTS中,该音频提示作为上下文学习(ICL),在不进行大规模参数更新或重训练的情况下引导模型的韵律与音色。为进一步提升生成质量并减少幻觉,我们引入一种基于ICL的在线强化学习策略,直接使用主观美学奖励优化自回归韵律模型,同时受连接时序分类(CTC)对齐约束以保证可懂性。综合人工感知评估表明,合成语音在自然度与表现力上均有显著提升,验证了基于ICL的在线强化学习方法的有效性。
原文摘要 · Abstract (English)
Conversational AI has made significant progress, yet generating expressive and controllable text-to-speech (TTS) remains challenging. Specifically, controlling fine-grained voice styles and emotions is notoriously difficult and typically requires massive amounts of heavily annotated training data. To overcome this data bottleneck, we present a scalable, data-efficient cascaded framework that pairs textual style tokens with human-curated, high-quality audio prompts. This approach enables single-shot adaptation to fine-grained speaking styles and character voices. In the context of TTS, this audio prompting acts as In-Context Learning (ICL), guiding the model's prosody and timbre without requiring massive parameter updates or large-scale retraining. To further enhance generation quality and mitigate hallucinations, we introduce a novel ICL-based online reinforcement learning (RL) strategy. This strategy directly optimizes the autoregressive prosody model using subjective aesthetic rewards while being constrained by Connectionist Temporal Classification (CTC) alignment to preserve intelligibility. Comprehensive human perception evaluations demonstrate significant improvements in both the naturalness and expressivity of the synthesized speech, establishing the efficacy of our ICL-based online RL approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。