arXiv:2606.15152cs.CL2026-06

测试智能体能否读懂视觉社交信号,发现其互动管理能力仍弱于局部表现。

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

论文配图:Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation
图 1 · 摘自论文原文
  • 构建多模态社交模拟基准,融合文本与视觉证据
  • 角色表达与冲突处理接近饱和,但互动调控仍困难
  • 适合研究视觉社交智能的AI系统评估

社交互动依赖语言与可见的社会信号,如面部表情、姿势、视线和情绪变化。然而现有社交智能体评测大多基于文本,很少检验多模态智能体是否能利用视觉线索指导互动。我们提出 extsc{AgentViSS},一个评估多模态社交模拟中视觉社交智能的基准。该基准包含240个场景、585个角色实例和2,340个角色-任务实例,结合对齐的文本-视觉证据、结构化角色档案以及四项角色级任务:表情任务、特征任务、互动调控任务和互动结果任务。在七种近期多模态大模型上,通过显式视觉输入与直接视觉输入两种方式评估,结果显示:角色特定的表情表达与冲突处理已接近饱和,而互动调控与视觉驱动的结果达成仍显著困难。代码开源于 https://github.com/JunsWan/AgentViSS,数据集可在 https://huggingface.co/datasets/JunsWan/AgentViSS 获取。

原文摘要 · Abstract (English)

Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are largely text-based and rarely test whether multimodal agents can use visual cues to guide interaction. We introduce \textsc{\benchmarkname{}}, a benchmark evaluating visual social intelligence in multimodal social simulation. It contains 240 scenarios, 585 role instances, and 2,340 role-task instances, combining aligned textual-visual evidence, structured role profiles, and four role-level tasks: expression task, characteristic task, interaction regulation task, and interaction outcome task. Evaluating seven recent MLLMs under verbalized-vision and direct-vision reveals a clear gap between local role enactment and interaction management: role-specific expression and conflict handling are near saturation, whereas interaction regulation and visually grounded outcome achievement remain substantially more difficult. The code is released at https://github.com/JunsWan/AgentViSS, and the dataset is available at https://huggingface.co/datasets/JunsWan/AgentViSS.

多模态社交智能视觉理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。