arXiv:2607.00946cs.SDcs.LG2026-07

对比两种语音生成模块的几何特性,揭示情绪控制效果差异。

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

论文配图:A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
图 1 · 摘自论文原文
  • 用线性探测和局部内在维数分析情绪表征几何结构
  • SLM模块情绪子空间清晰且说话人-情绪解耦好,CFM则存在耦合问题
  • 联合调控提升情绪强度但牺牲控制精度和语音质量

尽管已有研究探索混合文本到语音系统中的情绪控制,但这些模块的几何特性及其对可控性的意义仍不明确。本文首次对比了语音语言模型(SLM)与条件流匹配(CFM)模块作为激活调节点在混合情绪语音合成中的表现。通过线性探测和局部内在维数(LID)刻画情绪表示,并评估单点与联合调节在混合情绪合成中的效果。结果表明,SLM具有清晰、低维的情绪特定子空间,实现强说话人-情绪解耦;而CFM因说话人-情绪耦合导致跨说话人泛化能力差。联合调节虽增强情绪强度,但在分布内数据上降低比例控制精度并损害语音质量。研究为混合TTS系统的多点激活调节提供实践指导,并强调表示几何在可控语音生成中的关键作用。

原文摘要 · Abstract (English)

While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparative study of speech language model (SLM) and conditional flow-matching (CFM) modules as activation steering sites for mixed emotion speech synthesis. We first characterize emotion representations using linear probing and local intrinsic dimensionality (LID), and then evaluate single-site and joint steering for mixed-emotion synthesis. Our results show that SLM offers a clean, low-dimensional emotion-specific subspace with strong speaker--emotion disentanglement, while CFM exhibitspoor cross-speaker generalization due to speaker--emotion entanglement. Joint steering increases emotion intensity but degrades proportional control and speech quality on in-distribution data. These findings provide practical guidance for multi-site activation steering in hybrid TTS systems and highlight the importance of representation geometry in controllable speech generation.

语音合成情绪控制表示几何可调节生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。