arXiv:2509.19001eess.AScs.SD2025-09中稿 · ICASSP2026被引 3

通过分层解码实现指令驱动语音合成的精准控制

HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS

  • 将语音合成拆分为语义、风格、声学三阶段分层生成
  • 引入提示与内容偏好令牌,提升指令遵循率和自然度
  • 适合需要精细语音风格控制的研究与应用

基于大语言模型(LLM)的文生语音(TTS)模型已达到高度自然。然而,语音推理的精确控制仍具挑战性。尽管已有指令驱动的文生语音(Instruct-TTS)模型,但因单层文本指令与多层语音标记之间的模态鸿沟,仍缺乏细粒度控制能力。为此,我们提出HD-PPT框架,将语音合成转化为结构化的分层任务。为实现细粒度控制,引入新型语音编解码器,从复杂语音标记中提取由自动语音识别(ASR)和跨语言音频-文本预训练(CLAP)目标监督的提示偏好与内容偏好标记。为弥合这些标记的模态差异,提出分层解码策略:大语言模型按顺序生成语义、细粒度风格,最终生成完整声学表示。大量实验表明,该分层范式显著提升指令遵循度,并达到业界领先自然度,验证了该方法在精确可控语音合成中的有效性。音频样例见https://xxh333.github.io/。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging. Although instruction-based Text-to-Speech (Instruct-TTS) models are proposed, these models still lack fine-grained control due to the modality gap between single-level text instructions and multilevel speech tokens. To address this limitation, we propose HD-PPT, a framework that transforms speech synthesis into a structured, hierarchical task. To enable fine-grained control, we introduce a novel speech codec to extract distinct prompt-preference and content-preference tokens from the complex speech tokens, supervised by automatic speech recognition (ASR) and cross-lingual audio-text pre-training (CLAP) objectives. To bridge the modality gap of these tokens, we propose a hierarchical decoding strategy, where the LLM generates tokens in a structured order: first semantic, then fine-grained style, and finally complete acoustic representation. Extensive experiments demonstrate that this hierarchical paradigm significantly improves instruction adherence and achieves state-of-the-art naturalness, validating our approach for precise and controllable speech synthesis. Audio samples are available at https://xxh333.github.io/.

语音合成指令控制分层生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。