arXiv:2607.06461eess.AScs.CL2026-07

让大模型语音合成可精准操控每个词的发音特征。

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

论文配图:WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
图 1 · 摘自论文原文
  • 用绑定标记显式规划语音韵律,实现多维度控制。
  • 构建4.7小时双语标注数据集,支持五维声学属性调控。
  • 适合需要精细语音编辑的配音、有声书等场景使用。

尽管基于大语言模型(LLM)的文本转语音(TTS)系统已达到高度自然的效果,但主要依赖隐式的端到端生成范式,导致控制粒度粗糙。在有声书朗读和视频配音等需精确风格干预与严格时序对齐的场景中,无法显式操纵词级声学属性成为关键瓶颈。这一问题源于细粒度标注数据严重稀缺,以及将多维控制信号融入离散自回归生成的架构挑战。为此,本文提出统一框架以实现高精度词级控制。首先,构建了包含五维词级标注(持续时间、边界、能量、音高和语调)的大型双语数据集WordVoice-5A,总时长4.7小时,通过严谨的语言学引导流程生成。其次,提出WordVoice方法,将隐式生成转化为显式可控范式:在LLM中引入绑定标记机制,形成显式“声学规划”过程,支持自适应多任务韵律规划与灵活人工干预;同时,在标记到波形阶段引入细粒度声学调制模块,弥合离散标记与连续波形间的分辨率差距,确保词级属性严格对齐。大量实验表明,WordVoice在保持竞争性零样本合成稳定性的同时,实现了多维声学属性的优越且解耦的控制能力。代码与音频样例已公开于 https://xxh333.github.io/wordvoice-demo/。

原文摘要 · Abstract (English)

While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In scenarios demanding precise stylistic interventions and strict temporal alignment, such as audiobook narration and video dubbing, the inability to explicitly manipulate word-level acoustic attributes remains a critical bottleneck. This limitation is primarily amplified by the severe scarcity of fine-grained annotated datasets and the architectural challenge of integrating multi-dimensional control signals into discrete autoregressive generation. To address this, we propose a unified framework for highly precise word-level control. First, we construct WordVoice-5A, a massive 4.7k-hour bilingual dataset featuring five-dimensional word-level annotations (duration, boundary, energy, pitch and tone) developed through a rigorous linguistically-guided pipeline. Second, we introduce WordVoice to transform the implicit generation process into an explicit, highly controllable paradigm. Specifically, we introduce a bound-token mechanism within the LLM to formulate an explicit ``acoustic planning'' process, enabling adaptive multi-task prosodic planning and flexible manual intervention. Furthermore, we augment the token-to-waveform stage with a fine-grained acoustic modulation module, bridging the resolution gap to strictly align word-level attributes between highly compressed discrete tokens and continuous waveforms. Extensive experiments demonstrate that WordVoice achieves superior, decoupled control over multiple acoustic dimensions while maintaining competitive zero-shot synthesis stability. The code and audio samples are publicly available at https://xxh333.github.io/wordvoice-demo/.

语音合成大模型控制生成多维调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。