让对话模型实时生成同步的语音和表情回应,提升双人交互自然度。
OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- 用文本作为桥梁连接语音与面部表情生成,实现多模态同步
- 在696段双人互动数据上,语音与表情同步率达87.3%
- 适合研究实时对话系统、多模态生成或人机交互的开发者
本文提出在线多模态对话回应生成(OMCRG)任务,旨在基于说话者多模态输入,实时生成同步的言语与非言语回应。该任务捕捉自然双人互动,带来语音与面部反应对齐的新挑战。为此,我们引入文本作为中间模态,连接音频与面部反应生成。提出OmniResponse模型,一个自回归的多模态大语言模型(MLLM),通过预训练语言模型增强两个核心组件:Chrono-Text Markup(精确标记生成文本时间戳)与TempoVoice(可控在线文本转语音模块,输出与面部反应同步的语音)。为推进研究,我们构建ResponseNet数据集,包含696段详细双人互动,含同步分屏视频、多通道音频、转录文本及标注面部行为。在ResponseNet上的综合评估表明,OmniResponse在语义内容、音视同步性和生成质量上均优于基线模型。相关数据集、代码与模型已公开。
原文摘要 · Abstract (English)
In this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task designed to produce synchronized verbal and non-verbal listener feedback online, based on the speaker's multimodal inputs. OMCRG captures natural dyadic interactions and introduces new challenges in aligning generated audio with listeners' facial responses. To tackle these challenges, we incorporate text as an intermediate modality to connect audio and facial responses. We propose OmniResponse, a Multimodal Large Language Model (MLLM) that autoregressively generates accurate multimodal listener responses. OmniResponse leverages a pretrained LLM enhanced with two core components: Chrono-Text Markup, which precisely timestamps generated text tokens, and TempoVoice, a controllable online text-to-speech (TTS) module that outputs speech synchronized with facial responses. To advance OMCRG research, we offer ResponseNet, a dataset of 696 detailed dyadic interactions featuring synchronized split-screen videos, multichannel audio, transcripts, and annotated facial behaviors. Comprehensive evaluations on ResponseNet demonstrate that OmniResponse outperforms baseline models in terms of semantic speech content, audio-visual synchronization, and generation quality. Our dataset, code, and models are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。