用稀疏自编码器解析语音合成模型的内部表示并实现精准控制
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

- 通过稀疏自编码器捕捉文本与语音共享流中的语义特征
- 干预后笑声概率从0.02升至0.79,可切换说话人性别与语速
- 适合想理解或调控语音合成过程的研究者与开发者
语言模型日益成为文本到语音(TTS)系统的核心,但我们对其在文本与生成语音标记共用单一残差流时所构建的表征仍知之甚少。我们在CosyVoice3的语言模型主干上训练了BatchTopK稀疏自编码器,并提出一种模态感知的自动解释管道,用于标注每个特征的触发来源——文本前缀上下文、1秒语音片段,或两者兼有。恢复出的特征具有可解释性,涵盖音素、笑声、口音提示和说话人性别。通过稀疏自编码器潜在空间的操控表明,这些特征具有因果性而非仅描述性:定向干预使笑声概率从0.02提升至0.79,实现说话人性别反转,并控制语速同时保持内容不变。因此,SAE特征既可用于解释模型,也可作为TTS合成的可控方向。
原文摘要 · Abstract (English)
Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream. We train BatchTopK sparse autoencoders on the LM backbone of CosyVoice3 and introduce a modality-aware auto-interp pipeline that labels each feature from where it fires-text-prefix context, 1-second speech clips, or both. The recovered features are interpretable, spanning phonemes, laughter, accent prompts and speaker gender. Steering through the SAE latent space shows these features are causal rather than merely descriptive: targeted interventions raise laughter probability from 0.02 to 0.79, flip perceived speaker gender, and control speech rate while preserving spoken content. SAE features thus serve both as interpretability objects and as control directions for TTS synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。