arXiv:2504.14482cs.CLcs.SD2025-04中稿 · ICME 2025被引 5

用智能代理生成更自然的多人对话语音,解决数据少且单调的问题。

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue

  • 三类智能体协作写剧本、合成语音、评估改进,迭代优化对话质量。
  • 构建了多语言、多角色、多轮次的MultiTalk语音数据集,涵盖多样话题。
  • 适合研究语音合成、对话系统或需要定制化语音数据的研究者使用。

语音合成对人机交互至关重要,但现有数据集因人工标注成本高,存在角色单一、场景有限、情感表达不足等问题。为此,我们提出DialogueAgents——一种融合脚本撰写、语音合成与对话评价三类智能体的混合式语音合成框架。该框架基于多样化角色库,通过持续迭代优化对话脚本并生成语音,显著提升合成对话的情感表现力和副语言特征。基于此框架,我们构建了MultiTalk数据集,包含双语、多角色、多轮次对话,覆盖广泛主题。大量实验验证了框架的有效性及数据集的高质量。相关代码与数据已开源(https://github.com/uirlx/DialogueAgents),以推动先进语音合成模型与定制化数据生成研究。

原文摘要 · Abstract (English)

Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited character diversity, contextual scenarios, and emotional expressiveness. To address these issues, we propose DialogueAgents, a novel hybrid agent-based speech synthesis framework, which integrates three specialized agents -- a script writer, a speech synthesizer, and a dialogue critic -- to collaboratively generate dialogues. Grounded in a diverse character pool, the framework iteratively refines dialogue scripts and synthesizes speech based on speech review, boosting emotional expressiveness and paralinguistic features of the synthesized dialogues. Using DialogueAgent, we contribute MultiTalk, a bilingual, multi-party, multi-turn speech dialogue dataset covering diverse topics. Extensive experiments demonstrate the effectiveness of our framework and the high quality of the MultiTalk dataset. We release the dataset and code https://github.com/uirlx/DialogueAgents to facilitate future research on advanced speech synthesis models and customized data generation.

语音合成对话系统数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。