arXiv:2509.14627cs.HCcs.AI2025-09

让对话机器人生成有情绪、有风格的自然语音。

Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech

  • 基于对话情绪与回应风格生成语音描述,驱动自然语音合成。
  • 构建多模态对话数据集,提升语音情感表达能力。
  • 适合语音交互、虚拟助手等需要真实感语音的应用场景。

人类对话融合语言、语音和视觉线索,每种媒介提供互补信息。例如,语音传递文本无法完全表达的情绪或语调。当前多模态大模型主要关注从多元输入生成文本响应,对生成自然且吸引人的语音关注较少。本文提出一种类人对话代理,根据对话氛围和回应风格信息生成语音响应。为此,我们构建了一个专注于语音的多感官对话数据集(MultiSensory Conversation, MSenC),以支持代理生成自然语音。随后,我们设计了一种基于多模态大模型的框架,生成文本响应与语音描述,用于合成包含副语言信息的语音。实验结果表明,同时利用视觉与音频模态可有效生成富有吸引力的语音。源代码已公开于 https://github.com/kimtaesu24/MSenC。

原文摘要 · Abstract (English)

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating text responses from diverse inputs, less attention has been paid to generating natural and engaging speech. We propose a human-like agent that generates speech responses based on conversation mood and responsive style information. To achieve this, we build a novel MultiSensory Conversation dataset focused on speech to enable agents to generate natural speech. We then propose a multimodal LLM-based model for generating text responses and voice descriptions, which are used to generate speech covering paralinguistic information. Experimental results demonstrate the effectiveness of utilizing both visual and audio modalities in conversation to generate engaging speech. The source code is available in https://github.com/kimtaesu24/MSenC

语音生成多模态对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。