arXiv:2603.16859cs.AI2026-03被引 7

提出社交互动评测基准SocialOmni,评估多模态大模型在对话中的实时反应能力。

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

  • 构建三维度评测框架:识别说话人、判断插话时机、生成自然打断语句。
  • 包含2000个感知样本和209个严格约束的交互生成实例,测试模型鲁棒性。
  • 发现感知准确率与社交响应能力脱节,为模型优化提供关键信号。

多模态大语言模型(OLMs)通过原生融合音频、视觉与文本重新定义人机交互。然而,现有基准仍局限于静态、以准确率为导向的任务,未能评估对话中动态线索的社交互动能力。为此,我们提出SocialOmni,一个全面的评测基准,从三个核心维度量化模型的对话交互能力:(i) 发言者分离与识别(谁在说话),(ii) 插话时机控制(何时插话),(iii) 自然插话生成(如何表达打断)。SocialOmni包含2000个感知样本和209个质量可控的交互生成实例,具有严格的时空与上下文约束,并引入受控的音视频不一致场景以测试模型鲁棒性。我们对12个主流OLMs进行了评测,揭示了模型间社交互动能力的显著差异。分析进一步表明,模型的感知准确率与其生成恰当打断的能力存在明显解耦,说明仅依赖理解类指标无法全面刻画对话社交能力。更令人鼓舞的是,SocialOmni的诊断结果为未来OLMs弥合感知与交互之间的鸿沟提供了可操作的信号。

原文摘要 · Abstract (English)

Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a critical gap in assessing social interactivity, the fundamental capacity to navigate dynamic cues in natural dialogues. To this end, we propose SocialOmni, a comprehensive benchmark that operationalizes the evaluation of this conversational interactivity across three core dimensions: (i) speaker separation and identification (who is speaking), (ii) interruption timing control (when to interject), and (iii) natural interruption generation (how to phrase the interruption). SocialOmni features 2,000 perception samples and a quality-controlled diagnostic set of 209 interaction-generation instances with strict temporal and contextual constraints, complemented by controlled audio-visual inconsistency scenarios to test model robustness. We benchmarked 12 leading OLMs, which uncovers significant variance in their social-interaction capabilities across models. Furthermore, our analysis reveals a pronounced decoupling between a model's perceptual accuracy and its ability to generate contextually appropriate interruptions, indicating that understanding-centric metrics alone are insufficient to characterize conversational social competence. More encouragingly, these diagnostics from SocialOmni yield actionable signals for bridging the perception-interaction divide in future OLMs.

多模态对话系统评测基准社会互动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。