让AI在对话中实时理解多人社交互动,关键在提前预测发言和关注身体信号。
Towards Online Multi-Modal Social Interaction Understanding
- 通过粗到细的对话预测和社交感知视觉提示,实现在线推理。
- 在两个数据集上三项任务均达领先水平,显著优于基线模型。
- 适合研究人机交互、多模态理解的开发者与研究人员。
本文提出新问题Online-MMSI,要求模型仅基于历史信息完成多模态社交互动理解(MMSI)。给定一段记录视频和多方对话,智能助手需即时识别说话者指代对象,这对真实世界人机交互至关重要。由于无法获取未来对话上下文,人类和模型在从离线转为在线设置时性能大幅下降。为此,我们提出Online-MMSI-VLM框架,基于多模态大语言模型,核心创新包括:(1) 多方对话预测,以粗到细方式预判后续发言者及话语;(2) 社交感知视觉提示,通过边界框和人体关键点标注视频帧中的显著社交线索。该模型在两个数据集上的三项任务中均取得最优表现,显著超越基线,验证了Online-MMSI-VLM的有效性。项目页:https://sampson-lee.github.io/online-mmsi-project-page。
原文摘要 · Abstract (English)
In this paper, we introduce a new problem, Online-MMSI, where the model must perform multimodal social interaction understanding (MMSI) using only historical information. Given a recorded video and a multi-party dialogue, the AI assistant is required to immediately identify the speaker's referent, which is critical for real-world human-AI interaction. Without access to future conversational context, both humans and models experience substantial performance degradation when moving from offline to online settings. To tackle the challenges, we propose Online-MMSI-VLM, a novel framework based on multimodal large language models. The core innovations of our approach lie in two components: (1) multi-party conversation forecasting, which predicts upcoming speaker turns and utterances in a coarse-to-fine manner; and (2) socially-aware visual prompting, which highlights salient social cues in each video frame using bounding boxes and body keypoints. Our model achieves state-of-the-art results on three tasks across two datasets, significantly outperforming the baseline and demonstrating the effectiveness of Online-MMSI-VLM. Project page: https://sampson-lee.github.io/online-mmsi-project-page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。