arXiv:2604.08363cs.SD2026-04被引 3

用自然语言描述生成语音,支持单句和对话场景的统一建模。

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation

论文配图:CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
图 1 · 摘自论文原文
  • 基于标题条件的自回归框架,分层建模语音音色与表达。
  • 对话中实现逐轮动态属性控制,保持音色稳定且表达适配上下文。
  • 首个支持对话语音设计的统一方法,适合语音合成与人机交互研究者。

从自然语言描述生成语音是文本到语音多模态生成中的新兴任务,旨在不依赖参考音频的情况下合成具有目标音色和说话风格的语音。然而现有方法主要关注单句生成,对话场景的语音设计仍待探索。本文提出CapTalk,一种统一的标题条件文本-音频自回归框架,适用于单句和对话语音设计。在单句场景中使用句级描述,在对话中使用说话人级描述,并引入思维链控制序列以显式规划逐轮动态属性。为解决音色稳定性与上下文自适应表达之间的矛盾,提出分层变分条件模块,结合句级说话人编码器,实现音色复用同时保持表达对当前语句及对话上下文的适应性。构建了涵盖单句与对话场景的全面评估协议。实验表明,CapTalk在单句语音设计基准上达到最先进性能,并在多轮对话中展现出更优的表达可控性与情境恰当性。音频样本可访问:https://anonymous.4open.science/api/repo/Captalk-D601/file/index.html。

原文摘要 · Abstract (English)

Voice design from natural language descriptions is emerging as a new task in text-to-speech multimodal generation, aiming to synthesize speech with target timbre and speaking style without relying on reference audio. However, existing methods mainly focus on single-utterance generation, leaving conversational voice design largely unexplored. In this work, we extend voice design to dialogue, enabling better target speaker modeling and turn-level expressive control in natural conversational settings. We propose CapTalk, a unified caption-conditioned text-audio autoregressive framework for both single-utterance and dialogue voice design. CapTalk uses utterance-level captions for single-utterance voice design and speaker-level captions for dialogue speaker modeling, and further introduces a CoT control sequence in dialogue to explicitly plan turn-level dynamic attributes. To resolve the conflict between stable timbre preservation and context-adaptive expression, we propose a hierarchical variational conditioning module with an utterance-level speaker encoder to better balance stable timbre preservation and context-adaptive expression. This enables timbre reuse while keeping expression adaptive to the current utterance and, in dialogue, the surrounding context. We also build a comprehensive evaluation protocol for both single-utterance and dialogue settings. Experiments show that CapTalk achieves state-of-the-art performance on a single-utterance voice design benchmark and delivers better expression controllability and contextual appropriateness in multi-turn dialogue. Audio samples are available at: https://anonymous.4open.science/api/repo/Captalk-D601/file/index.html.

语音生成对话系统多模态文本到语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。