用稀疏自编码器挖掘语音合成中的可解释情感特征
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

- 通过稀疏自编码器识别文本到语音模型中分散的情感特征
- 仅干预少量稀疏特征即可实现可解释的情感控制,效果优于全局调节
- 发现不同特征对应特定音高属性,适合需要可控情感表达的研究者
将大语言模型(LLM)融入文本到语音(TTS)系统提升了语音表现力,但可解释的情感控制仍具挑战。现有方法多依赖外部条件或全局激活调节,难以揭示内部表征机制。本文利用稀疏自编码器(SAEs)分析基于LLM的TTS模型语义隐藏状态中的情感变化,发现情感差异分布在多个稀疏潜在特征中,而干预其中一小部分即可实现可解释的情感调控。基于此,我们提出无需修改主干参数的特征级干预框架,支持双向情感诱导与抑制。进一步发现,不同潜在特征关联特定声学属性(如音高),表明情感表达源于多特征协同而非单一全局偏移。实验表明,该方法在情感诱导与抑制上达到或超过全局调节及现有TTS基线性能。
原文摘要 · Abstract (English)
Integrating large language models (LLMs) into text-to-speech (TTS) systems has improved speech expressiveness, yet interpretable emotional control remains challenging. Existing approaches primarily rely on external conditioning or global activation steering, offering limited insight into the internal representations underlying emotional control. In this work, we analyze emotion-related variation in the semantic hidden states of LLM-based TTS models using sparse autoencoders (SAEs) to identify sparse latent features. Our analysis shows that emotional variation is distributed across multiple sparse latent features, while intervening on a small subset enables interpretable emotion control. Building on this observation, we introduce a feature-level intervention framework for bidirectional emotion induction and suppression without modifying backbone parameters. We further show that distinct latent features are associated with specific acoustic attributes (e.g., pitch), suggesting that emotional expression arises from coordinated latent contributions rather than a single global shift. Empirically, steering these sparse latent features achieves comparable or superior emotion induction and suppression performance relative to global steering and existing TTS baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。