arXiv:2606.20532cs.AI2026-06

解析自然语言如何控制语音风格,揭示关键词对声音的动态影响

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

论文配图:How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
图 1 · 摘自论文原文
  • 用跨注意力归因分析语音扩散模型中每个词的作用
  • 发现风格词影响音高与能量,且在早期深层最显著
  • 适合想理解语音生成机制或优化可控性的研究者

风格化文本到语音系统利用自然语言控制语音特征,但单个词语如何影响声学输出仍不明确。我们首次将DAAM框架适配至语音领域,提出针对语音扩散模型的跨注意力归因方法,并应用于CapSpeech-TTS。该方法在25层和24个ODE步骤中提取每个词的热力图。分析了3600组(风格描述,文本原文)组合,每组由120个风格描述分别控制30个文本生成。结果表明:(1) 风格词的时间方差低于内容/功能词,证实全局调节;(2) 风格注意力与基频(F0)和能量相关;(3) 风格调节在早期步骤和深层网络最强;(4) 第17层注意力熵最低,与风格重要性峰值重合,表明该阶段网络选择性最强。这是首个研究自然语言如何影响语音扩散模型中跨注意力的论文。

原文摘要 · Abstract (English)

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeech-TTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models

语音生成扩散模型注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。