让聊天机器人同时具备看和听的能力,实现自然动态的多人多轮对话。
Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
- 构建新数据集M³C,支持多轮、多方、视听融合的复杂对话
- 模型可长期记忆并实时处理视觉与语音输入,保持对话连贯性
- 适合研究多模态对话系统或智能助手的开发者与研究人员
随着聊天机器人向类人化、现实世界交互演进,多模态仍为研究热点。现有工作多聚焦图像任务(如视觉对话),侧重‘视觉’而忽视‘听觉’;且多限于静态交互,难以实现自然动态参与。此外,虽已有针对多方、多轮对话的研究,但任务约束阻碍其在真实场景中无缝集成。为此,本文提出赋予聊天机器人‘眼耳并用’能力,构建全新多模态多轮多方对话数据集M³C,设计支持多模态记忆检索的新模型。该模型在M³C上训练后,可在复杂现实场景中与多位说话者进行长时程、动态交互,有效处理视听信息以实现恰当响应。人工评估显示模型在保持对话连贯性与动态性方面表现优异,展现出先进多模态对话代理的潜力。
原文摘要 · Abstract (English)
As chatbots continue to evolve toward human-like, real-world, interactions, multimodality remains an active area of research and exploration. So far, efforts to integrate multimodality into chatbots have primarily focused on image-centric tasks, such as visual dialogue and image-based instructions, placing emphasis on the "eyes" of human perception while neglecting the "ears", namely auditory aspects. Moreover, these studies often center around static interactions that focus on discussing the modality rather than naturally incorporating it into the conversation, which limits the richness of simultaneous, dynamic engagement. Furthermore, while multimodality has been explored in multi-party and multi-session conversations, task-specific constraints have hindered its seamless integration into dynamic, natural conversations. To address these challenges, this study aims to equip chatbots with "eyes and ears" capable of more immersive interactions with humans. As part of this effort, we introduce a new multimodal conversation dataset, Multimodal Multi-Session Multi-Party Conversation ($M^3C$), and propose a novel multimodal conversation model featuring multimodal memory retrieval. Our model, trained on the $M^3C$, demonstrates the ability to seamlessly engage in long-term conversations with multiple speakers in complex, real-world-like settings, effectively processing visual and auditory inputs to understand and respond appropriately. Human evaluations highlight the model's strong performance in maintaining coherent and dynamic interactions, demonstrating its potential for advanced multimodal conversational agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。