arXiv:2506.00958cs.AIcs.CL2025-06ACL被引 7

构建大规模多模态数据集,让对话模型学会理解并生成非语言表达。

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues

  • 基于视频对齐的文本与表情/动作标注,构建大尺度多模态数据集VENUS。
  • 训练统一框架模型MARS,实现文本与非语言信号的联合生成。
  • 适合研究多模态对话、情感计算及虚拟角色交互的开发者和研究者。

非语言交流是人类互动的核心,手势、面部表情和肢体语言传递着意图与情感的关键信息。然而现有大型语言模型(LLMs)难以有效融入这些非语言元素,限制了其创造沉浸式对话体验的能力。本文提出MARS,一种能够理解并生成非语言线索的多模态语言模型,以填补这一空白。核心创新在于VENUS——一个包含时间对齐的视频、文本、面部表情和身体语言标注的大规模数据集。基于VENUS,我们采用下一个词预测目标训练MARS,将文本与向量量化后的非语言表示结合,在统一框架中实现多模态理解与生成。通过对VENUS数据集的多维度分析,验证了其显著规模与高有效性。定量与定性结果表明,MARS能根据对话输入成功生成对应的文字与非语言表达。

原文摘要 · Abstract (English)

Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorporate these nonverbal elements, limiting their capacity to create fully immersive conversational experiences. We introduce MARS, a multimodal language model designed to understand and generate nonverbal cues alongside text, bridging this gap in conversational AI. Our key innovation is VENUS, a large-scale dataset comprising annotated videos with time-aligned text, facial expressions, and body language. Leveraging VENUS, we train MARS with a next-token prediction objective, combining text with vector-quantized nonverbal representations to achieve multimodal understanding and generation within a unified framework. Based on various analyses of the VENUS datasets, we validate its substantial scale and high effectiveness. Our quantitative and qualitative results demonstrate that MARS successfully generates text and nonverbal languages, corresponding to conversational input.

多模态对话系统非语言表达视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。