arXiv:2509.26514cs.CL2025-09

用LLM当指挥官,让语音合成更听话、更可控。

BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs

  • LLM解析指令生成语音特征计划,独立TTS按计划发声
  • 零样本跨语言控制成功,未见语言也能精准应用声调等特征
  • 适合需要高精度语音控制的场景,如虚拟角色、多语言语音助手

大型语言模型(LLMs)正重塑多模态系统,语音合成是其重要应用之一。然而,现有方法常未能充分利用这些模型的语言智能,尤其忽视其强大的指令遵循能力,限制了文本到语音(TTS)的可控性。为此,我们提出一种受“操作主义”启发的新范式,将指令理解与语音生成解耦。我们构建BatonVoice框架:由一个LLM充当“指挥官”,理解用户指令并生成包含显式声学特征(如音高、能量)的文本“计划”;再由独立的TTS模型(“乐团”)根据该计划生成语音。为实现此架构,我们开发了专用于该任务的BatonTTS模型。实验表明,BatonVoice在可控及情感语音合成上表现优异,优于多个开源与闭源基线模型。尤为突出的是,该方法实现了显著的零样本跨语言泛化能力,能准确将特征控制能力应用于训练阶段未见的语言。这表明,将语音对象化为文本声学特征,可更有效地激发LLMs的语言智能。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) is reshaping multimodel models, with speech synthesis being a prominent application. However, existing approaches often underutilize the linguistic intelligence of these models, typically failing to leverage their powerful instruction-following capabilities. This limitation hinders the model's ability to follow text instructions for controllable Text-to-Speech~(TTS). To address this, we propose a new paradigm inspired by ``operationalism'' that decouples instruction understanding from speech generation. We introduce BatonVoice, a framework where an LLM acts as a ``conductor'', understanding user instructions and generating a textual ``plan'' -- explicit vocal features (e.g., pitch, energy). A separate TTS model, the ``orchestra'', then generates the speech from these features. To realize this component, we develop BatonTTS, a TTS model trained specifically for this task. Our experiments demonstrate that BatonVoice achieves strong performance in controllable and emotional speech synthesis, outperforming strong open- and closed-source baselines. Notably, our approach enables remarkable zero-shot cross-lingual generalization, accurately applying feature control abilities to languages unseen during post-training. This demonstrates that objectifying speech into textual vocal features can more effectively unlock the linguistic intelligence of LLMs.

语音合成大模型可控生成跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。