让语音合成更自然地表达情绪冲突,提升情感识别准确率
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
- 根据文本与语音情绪不一致程度动态调整引导强度
- 情绪识别准确率提升12%,主观评分提高10%
- 适合需要精准情感控制的语音合成应用
尽管文本到语音(TTS)系统可通过自然语言指令实现情感控制,但当目标情感与文本语义冲突时,表现力、自然度和语音质量会下降。本文提出一种基于跨模态一致性引导的无分类器引导方法(CCG-CFG),其引导强度随文本情绪与语音情绪的不一致程度动态调整,并将丢弃条件替换为文本情绪。同时采用硬样本挖掘策略对引导信号进行蒸馏,增强模型的情绪对齐能力。在五个情感语料库和两个TTS基准上的评估表明,该方法应用于CosyVoice2后,情绪识别准确率最高提升12%绝对值,主观评分相对提升10%,优于HierSpeech++、Qwen3-TTS及原始CosyVoice2,且保持可懂性、自然度和高质量语音输出。
原文摘要 · Abstract (English)
While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion conflicts with the textual semantics. We propose a Cross-modal Consistency Guided Classifier-Free Guidance (CCG-CFG) method with dynamic scales based on the degree of inconsistency between the text emotion and the explicit speech emotion, replacing the dropout condition with the text emotion. We also distill the CCG-CFG guidance signal using a hard-sample mining strategy, improving the TTS model's emotional alignment capability. Evaluations on five emotional corpora and two TTS benchmarks show that our approaches applied to CosyVoice2 achieve up to a 12% absolute improvement in emotion-recognition accuracy and a 10% relative improvement in subjective scores, outperforming baselines including HierSpeech++, Qwen3-TTS, and original CosyVoice2, while preserving intelligibility, naturalness, and high speech quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。