arXiv:2609.02623cs.SDcs.CL2026-09

用伪三元组让语音合成按指令改风格,还能保原音色。

Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

论文配图:Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
图 1 · 摘自论文原文
  • 用可控音色的语音模型生成风格变化,再让大模型写自然语言指令
  • 仅用伪数据就能稳定保持说话人特征并实现方向性修改
  • 结合真实数据可更好对齐指令,适合语音定制与演播场景

语音演员常在不改变台词的前提下,根据表演指令重读同一剧本。我们研究此场景下的方向跟随语音合成(direction-following TTS),即系统在保留说话人身份和语义内容的前提下,生成符合给定方向的新语音。核心挑战在于缺乏此类相对变化的训练数据。为此,我们提出一种可扩展的伪三元组构建流水线,自动生成(参考语音,方向文本,修改后语音)三元组。该方法利用可控音色的语音合成模型生成受控风格变化,并通过大模型从估算的音色差异中生成自然语言方向描述。实验表明,仅使用伪三元组即可实现稳定的说话人保持式修改;结合真实数据能进一步提升方向对齐效果,同时维持说话人相似性。音频示例可在演示页查看:https://ntt-hilab-gensp.github.io/IS2026pseudo/

原文摘要 · Abstract (English)

Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page https://ntt-hilab-gensp.github.io/IS2026pseudo/

语音合成风格控制伪数据指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。