arXiv:2505.19978cs.CLcs.SD2025-05被引 9

构建首个大规模多轮情感对话数据集,支持语音与情绪同步。

DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset

  • 用9个大模型生成初稿,再经人工+AI筛选高质量对话。
  • 40,150条对话覆盖41个领域,含20种情绪且连贯推进。
  • 首次实现情绪一致的语音合成,适合多模态对话研究。

近期对话人工智能在单轮回复上表现优异,但多轮对话仍对最先进语言模型构成挑战。现有对话数据集在情感范围、领域多样性、对话轮次深度方面受限,且多为纯文本,阻碍了跨模态类人对话系统的发展。为此,我们推出DeepDialogue,一个大规模多模态数据集,包含40,150条高质量多轮对话,覆盖41个领域,涵盖20种不同情绪并具有连贯的情感演进。我们采用9个参数量从4B到72B的语言模型生成65,600条初始对话,再通过人工标注与基于LLM的质量过滤进行评估。结果揭示:小模型在6轮后难以保持连贯性;具体领域(如“汽车”、“旅行”)比抽象领域(如“哲学”)产生更有意义对话;跨模型交互对话比同模型对话更连贯。DeepDialogue的核心贡献在于其语音组件——为全部40,150条对话合成情感一致的声音,成为首个大规模开源、忠实保留情感上下文的多轮对话多模态数据集。

原文摘要 · Abstract (English)

Recent advances in conversational AI have demonstrated impressive capabilities in single-turn responses, yet multi-turn dialogues remain challenging for even the most sophisticated language models. Current dialogue datasets are limited in their emotional range, domain diversity, turn depth, and are predominantly text-only, hindering progress in developing more human-like conversational systems across modalities. To address these limitations, we present DeepDialogue, a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. Our approach pairs 9 different language models (4B-72B parameters) to generate 65,600 initial conversations, which we then evaluate through a combination of human annotation and LLM-based quality filtering. The resulting dataset reveals fundamental insights: smaller models fail to maintain coherence beyond 6 dialogue turns; concrete domains (e.g., "cars," "travel") yield more meaningful conversations than abstract ones (e.g., "philosophy"); and cross-model interactions produce more coherent dialogues than same-model conversations. A key contribution of DeepDialogue is its speech component, where we synthesize emotion-consistent voices for all 40,150 dialogues, creating the first large-scale open-source multimodal dialogue dataset that faithfully preserves emotional context across multi-turn conversations.

对话系统多模态情感识别语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。