arXiv:2510.13105cs.CV2025-10被引 3

评测大模型在社交场景中主动干预的能力,提出新数据集与检测方法。

EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception

  • 构建1.35万对视角化社交视频问答数据集,用于评估模型干预时机判断能力。
  • 现有大模型干预时机识别准确率仅14.4%(Gemini 2.5 Pro),表现不佳。
  • 提出EgoSoD方法,融合音视频多模态信息,提升主动干预与社交理解性能。

随着AR/VR技术融入日常生活,亟需能从第一人称视角理解人类社交动态的AI。然而当前大语言模型普遍缺乏社会意识,难以判断何时应作为助手介入,常产生不合时宜的回应,干扰自然对话并影响用户专注。为此,我们提出EgoSocial,一个包含13,500对社交视频-问题对的大规模第一人称视角数据集,专门用于评估模型在社交互动感知中的主动干预能力。我们对当前多模态大模型(OLLMs)进行了深入分析,发现其在识别多样社交上下文线索方面仍存在明显不足。实验表明,现有模型在干预时机识别上的准确率仅为14.4%(Gemini 2.5 Pro)。为此,我们提出EgoSoD(EgoSocial Detection),一种端到端方法,通过构建社会思维图,动态整合音频、视觉等多模态上下文线索,精准识别干预时机与社交互动。EgoSoD使Phi-4在干预时机任务上提升45.6%,Gemini 2.5 Pro提升9.9%;在整体社交互动任务上,分别提升20.4%和6.9%。相关数据集与代码即将开源。

原文摘要 · Abstract (English)

As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs often lack the social awareness to discern when to intervene as AI assistant. This leads to constant, socially unaware responses that may disrupt natural conversation and negatively impact user focus. To address these limitations, we introduce EgoSocial, a large-scale egocentric dataset with 13,500 social video-question pairs, specifically designed to benchmark intervention in social interaction perception. We also present an in-depth analysis of current omnimodal LLMs (OLLMs) to assess their effectiveness in detecting diverse social contextual cues. Experiments show that OLLMs still struggle to detect the intervention timing (14.4% for Gemini 2.5 Pro). We also propose EgoSoD (EgoSocial Detection), an end-to-end method for robustly discerning social dynamics. Informed by our OLLM analysis, EgoSoD integrates multimodal contextual cues (e.g., audio and visual cues) into a social thinking graph, dynamically modeling participants and interactions. Our method proactively detects intervention timing and social interactions, precisely determining when to intervene. Our EgoSoD improves Phi-4 by 45.6% and Gemini 2.5 Pro by 9.9% on Intervention Timing performance, and improves Phi-4 by 20.4% and Gemini 2.5 Pro by 6.9% on overall Social Interaction performance. We will release the dataset and code soon.

多模态社交智能主动干预第一人称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。