arXiv:2608.16053cs.CLeess.AS2026-08

让对话生成更自然:内容、时机、语音分开处理,真实感更强。

Agentic-DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech

论文配图:Agentic-DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech
图 1 · 摘自论文原文
  • 用双模型实时互听互答,让对话节奏自然涌现
  • 生成的对话在重叠、打断等动态上更接近真人对谈
  • 适合需要高真实感语音交互的医疗、客服场景

合成对话语音已成为开发和评估对话系统的重要资源。然而,现有对话合成流程通常先生成对话内容,再通过手工标记或时间规则插入打断、重叠和回应词,导致对话节奏人为预设而非互动驱动。本文提出 Agentic-DuplexGen 框架,明确解耦内容、时机与音质。首先由大语言模型生成对话脚本,随后两个全双工对话模型实时互听互答执行脚本。这使得对话时机自然生成,同时保留脚本内容。最后,一个高保真语音合成模型重新渲染交互过程,不改变原有时间结构。为验证该框架,我们构建了一个患者-医生对话语音语料库,包含构建时标注的词时间戳、说话人活动、重叠区域及交互事件。实验结果表明,该框架生成的对话动态比传统拼接式合成更接近真实对话。

原文摘要 · Abstract (English)

Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert interruptions, overlap, and backchannels using handcrafted markers or timing rules, making conversational timing prescribed rather than interaction-driven. We present Agentic-DuplexGen, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics. An LLM first generates the dialogue script, and then two full-duplex conversational models perform the script while listening to each other in real time. This allows conversational timing to emerge naturally while preserving the scripted content. Finally, a high-fidelity text-to-speech model re-renders the interaction without altering its timing. As a demonstration of the proposed framework, we construct a patient--clinician conversational speech corpus with construction-time annotations, including word timestamps, speaker activity, overlap regions, and interaction events. Experimental results show that the proposed framework produces conversational dynamics closer to real dialogue than conventional stitching-based synthesis.

对话生成语音合成自然对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。