arXiv:2601.22873eess.AScs.AI2026-01中稿 · ICASSP 2026被引 4

用轻量级向量调控语音合成情绪,让情感表达更自然可控。

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

  • 通过学习情绪偏移向量,动态调节输出特征空间中的情感表现。
  • 仅1000万参数,性能优于零样本和全微调基线。
  • 适合需要精准控制情感强度的语音合成应用。

在文本到语音(TTS)合成中,精确且可控的情感表达对生成自然、符合语境的语音至关重要。然而,许多情感感知型TTS系统,包括基于大语言模型(LLM)的设计,依赖于固定的语气嵌入或外部引导,限制了其对特定情绪潜在特征的建模能力。为此,我们提出EmoShift,一种轻量级激活调控框架,包含一个EmoSteer层,该层在输出嵌入空间中为每种目标情绪学习一个调控向量,以捕捉其潜在偏移,并在不同话语和类别间保持稳定、恰当的情感表达。该方法仅需1000万可训练参数,不到全微调的1/30,在客观与主观评估中均优于零样本及全微调基线,显著提升情感表现力的同时保持语音自然度与说话人相似性。进一步分析证实了EmoSteer层的有效性,并揭示其在可控情感强度方面的潜力。

原文摘要 · Abstract (English)

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based designs, rely on scaling fixed emotion embeddings or external guidance, limiting their ability to model emotion-specific latent characteristics. To address this gap, we present EmoShift, a lightweight activation-steering framework incorporating a EmoSteer layer, which learns a steering vector for each target emotion in the output embedding space to capture its latent offset and maintain stable, appropriate expression across utterances and categories. With only 10M trainable parameters,less than 1/30 of full fine-tuning, EmoShift outperforms zero-shot and fully fine-tuned baselines in objective and subjective evaluations, enhancing emotional expressiveness while preserving naturalness and speaker similarity. Further analysis confirms the proposed EmoSteer layer's effectiveness and reveals its potential for controllable emotional intensity in speech synthesis.

语音合成情感控制轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。