arXiv:2510.08373eess.AScs.SD2025-10被引 2

用大模型与流匹配生成自然对话语音,支持双人交互与跨语言

DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching

  • 结合大语言模型与分块流匹配,构建双轨对话语音生成架构
  • 在多轮对话中实现自然的发言交替与重叠语音,提升语境连贯性
  • 支持中英双语及跨语言合成,适合语音助手、虚拟角色开发

近年来,基于大语言模型(LLM)的文本转语音(TTS)合成显著提升了表达力和自然度。然而,生成类人、互动式的对话语音仍具挑战,现有系统受限于双轨数据稀缺,难以实现自然度、语境连贯性及互动动态(如换言、重叠说话、说话人一致性)等多轮对话特性。为此,我们提出DialoSpeech,一种融合大语言模型与分块流匹配的双轨架构,实现富有表现力的人类级对话语音合成。该方法生成具有连贯说话人轮次与自然重叠的多轮对话,支持中英文及跨语言语音合成。我们设计了数据处理流程以构建双轨对话数据集,支持可扩展训练与实验验证。实验表明,该模型优于基线方法,为生成类人对话语音提供有效解决方案。音频样例见 https://tiamojames.github.io/DialoSpeech

原文摘要 · Abstract (English)

Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However, generating human-like, interactive dialogue speech remains challenging. Current systems face limitations due to the scarcity of dual-track data and difficulties in achieving naturalness, contextual coherence, and interactional dynamics, such as turn-taking, overlapping speech, and speaker consistency, in multi-turn conversations. To address these challenges, we propose DialoSpeech, a dual-track architecture combining a large language model with Chunked Flow Matching for expressive, human-like dialogue speech synthesis. DialoSpeech generates natural multi-turn conversations with coherent speaker turns and natural overlaps, supporting both Chinese and English and cross-lingual speech synthesis. We introduce a data processing pipeline to construct dual-track dialogue datasets, facilitating scalable training and experimental validation. Experiments show that our model outperforms baselines, offering a solution for generating human-like spoken dialogues. Audio samples are available at https://tiamojames.github.io/DialoSpeech

对话生成语音合成大模型双轨对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。