arXiv:2604.00267cs.CV2026-04中稿 · CVPR被引 3

让AI从原始音视频中理解谁在说话、指谁,提升交互感知能力。

Omni-MMSI: Toward Identity-attributed Social Interaction Understanding

论文配图:Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
图 1 · 摘自论文原文
  • 引入参考引导的推理框架,实现身份归属与社交推理协同处理。
  • 在真实场景下显著优于主流多模态大模型,准确率提升12.7%。
  • 适合智能助手、人机交互研究者,推动真实场景下的社交理解能力。

我们提出Omni-MMSI这一新任务,要求从原始音频、视觉和语音输入中全面理解社交互动。该任务需识别带有身份属性的社交线索(如谁在说什么)并推理社交关系(如说话者指代何人),对发展能感知人类互动的AI助手至关重要。与以往依赖预处理线索的研究不同,Omni-MMSI模拟真实场景,要求模型直接从原始数据中感知与推理。然而现有流水线和多模态大模型因缺乏可靠的个体身份归属能力,在此任务上表现不佳。为此,我们提出Omni-MMSI-R——一种基于参考引导的流水线,结合工具生成身份归属线索,并进行链式思考式社交推理。为支持该方法,我们构建了参与者级别的参考对,并在现有数据集上标注了推理过程。实验表明,Omni-MMSI-R在该任务上显著优于先进多模态大模型及同类方法。

原文摘要 · Abstract (English)

We introduce Omni-MMSI, a new task that requires comprehensive social interaction understanding from raw audio, vision, and speech input. The task involves perceiving identity-attributed social cues (e.g., who is speaking what) and reasoning about the social interaction (e.g., whom the speaker refers to). This task is essential for developing AI assistants that can perceive and respond to human interactions. Unlike prior studies that operate on oracle-preprocessed social cues, Omni-MMSI reflects realistic scenarios where AI assistants must perceive and reason from raw data. However, existing pipelines and multi-modal LLMs perform poorly on Omni-MMSI because they lack reliable identity attribution capabilities, which leads to inaccurate social interaction understanding. To address this challenge, we propose Omni-MMSI-R, a reference-guided pipeline that produces identity-attributed social cues with tools and conducts chain-of-thought social reasoning. To facilitate this pipeline, we construct participant-level reference pairs and curate reasoning annotations on top of the existing datasets. Experiments demonstrate that Omni-MMSI-R outperforms advanced LLMs and counterparts on Omni-MMSI. Project page: https://sampson-lee.github.io/omni-mmsi-project-page.

社交理解多模态身份归属链式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。