arXiv:2508.03543cs.SDcs.AI2025-08被引 18

无需训练即可精细调控语音情感,支持任意情绪转换与插值。

EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering

  • 通过调节流匹配模型内部激活值实现情感控制
  • 在多款预训练模型上实现连续且可解释的情绪调节
  • 适合需要灵活情感编辑的语音合成应用

近年来,文本到语音(TTS)技术取得了显著进展。然而,现有TTS系统大多仅提供粗粒度、僵化的感情控制,通常依赖离散的情感标签或精心设计的详细情感提示,导致精细情感操作难以实现或不稳定。这些模型还需大量高质量数据进行训练。为解决上述问题,我们提出EmoSteer-TTS,一种全新的无需训练的方法,通过激活调节实现细粒度语音情感控制(转换、插值、消除)。我们首先实证发现,修改基于流匹配的TTS模型中的部分内部激活值可有效改变合成语音的情感基调。基于此,我们开发了一套无需训练、高效的算法,包括激活提取、情感标记搜索和推理时调节,可无缝集成至多种预训练模型(如F5-TTS、CosyVoice2和E2-TTS)。此外,为生成有效的调节向量,我们构建了一个包含多样化说话人的精选情感语音数据集。大量实验表明,EmoSteer-TTS实现了细粒度、可解释且连续的情感控制,优于当前最优方法(SOTA)。据我们所知,这是首个实现无需训练且连续细粒度情感控制的TTS方法。演示样本见https://emosteer-tts-demo.pages.dev/。

原文摘要 · Abstract (English)

Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emotional text prompt, making fine-grained emotion manipulation either inaccessible or unstable. These models also require extensive, high-quality datasets for training. To address these limitations, we propose EmoSteer-TTS, a novel training-free approach, to achieve fine-grained speech emotion control (conversion, interpolation, erasure) by activation steering. We first empirically observe that modifying a subset of the internal activations within a flow matching-based TTS model can effectively alter the emotional tone of synthesized speech. Building on this insight, we then develop a training-free and efficient algorithm, including activation extraction, emotional token searching, and inference-time steering, which can be seamlessly integrated into a wide range of pretrained models (e.g., F5-TTS, CosyVoice2, and E2-TTS). In addition, to derive effective steering vectors, we construct a curated emotional speech dataset with diverse speakers. Extensive experiments demonstrate that EmoSteer-TTS enables fine-grained, interpretable, and continuous control over speech emotion, outperforming the state-of-the-art (SOTA). To the best of our knowledge, this is the first method that achieves training-free and continuous fine-grained emotion control in TTS. Demo samples are available at https://emosteer-tts-demo.pages.dev/.

语音合成情感控制无训练激活调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。