arXiv:2602.03420cs.SDcs.LG2026-02被引 4

让语音合成同时表达多种情绪,更像真人说话。

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

  • 通过激活向量控制情感,实现多情绪混合合成
  • 发现语言模块主导情感韵律生成,非流匹配模块
  • 支持文本与情绪不一致的自然表达,适合影视配音

人类语音中的情感表达复杂且具有组合性,常包含多个甚至相互冲突的情感线索,可能与语言内容不一致。而现有表现力强的文语转换系统通常只施加单一情感级别,忽略了情感多样性,抑制了混合或语义-情绪不符的表达。尽管通过潜在方向向量进行激活调控提供了一种可行方案,但尚不清楚情感表征在文语转换中是否可线性调控,调控应作用于何种混合架构,以及如何评估此类复杂情感行为。本文首次系统分析了混合文语转换模型中激活调控的情感控制能力,提出一种定量、可控的调控框架及多评价者评估协议,实现了可组合的混合情感合成与可靠的语义-情绪不一致合成。结果首次表明,情感韵律和表达变异性主要由文语转换的语言模块生成,而非流匹配模块,并提供一种轻量级调控方法,用于生成自然、类人的情感语音。

原文摘要 · Abstract (English)

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech systems enforce a single utterance-level emotion, collapsing affective diversity and suppressing mixed or text-emotion-misaligned expression. While activation steering via latent direction vectors offers a promising solution, it remains unclear whether emotion representations are linearly steerable in TTS, where steering should be applied within hybrid TTS architectures, and how such complex emotion behaviors should be evaluated. This paper presents the first systematic analysis of activation steering for emotional control in hybrid TTS models, introducing a quantitative, controllable steering framework, and multi-rater evaluation protocols that enable composable mixed-emotion synthesis and reliable text-emotion mismatch synthesis. Our results demonstrate, for the first time, that emotional prosody and expressive variability are primarily synthesized by the TTS language module instead of the flow-matching module, and also provide a lightweight steering approach for generating natural, human-like emotional speech.

语音合成情感控制混合模型激活调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。