arXiv:2511.00261cs.CVcs.HC2025-11

用体育场景测试模型推断隐藏球位置的能力,发现人类远超模型。

Spot The Ball: A Benchmark for Visual Social Inference

  • 以球类运动图像为背景,让模型根据他人姿态眼神推断被移除的球
  • 人类准确率20%-34%,模型最高仅17%,差距达2-3倍
  • 模型依赖位置猜测,人类善用眼神与姿态等社交线索

人类在视觉社交推理方面表现优异,能通过他人注视方向、姿势和朝向等细微行为线索推断场景中隐藏的信息。这种能力对日常社会推理至关重要,也是构建类人智能体的关键。我们提出Spot The Ball基准,利用体育场景评估视觉语言模型(VLMs)在视觉社交推理方面的能力,任务是定位被移除的足球、篮球和排球。我们构建了一个经过筛选的评估集,包含人类基线,并设计了可扩展的生成流水线以生成更多测试样本。我们采用三种提示策略,评估了四种顶尖VLM(Gemini、GPT、LLaMA、Qwen),结果显示,在所有运动项目中,人类准确率始终为模型的2至3倍(人类20-34%,模型≤17%)。分析表明,模型主要依赖空间位置的表面启发式规则(如猜中心或靠近球员),而人类则善于利用注视方向和身体姿态等社交线索。这些发现揭示了人类与模型在视觉社交推理上的持续差距,强调需要设计能显式编码结构化行为线索的架构,以实现更稳健、类人的推理能力。

原文摘要 · Abstract (English)

Humans excel at visual social inference, the ability to infer hidden elements of a scene from subtle behavioral cues such as other people's gaze, pose, and orientation. This ability drives everyday social reasoning in humans and is critical for developing more human-like AI agents. We introduce Spot The Ball, a challenging benchmark for evaluating visual social inference in vision-language models (VLMs) using sports as a test domain. The task is to localize a removed sports ball from soccer, basketball, and volleyball images. We present a curated evaluation set with human baselines and a scalable pipeline for generating additional test items. We evaluate four state-of-the-art VLMs (Gemini, GPT, LLaMA, Qwen) using three prompting strategies, finding that humans are consistently two to three times more accurate (20-34%) than models ($\leq$ 17%) across all sports. Our analyses show that models rely on superficial spatial heuristics--such as guessing near the image center or nearby players--while humans leverage social cues like gaze direction and body pose. These findings reveal a persistent human-model gap in visual social reasoning and underscore the need for architectures that explicitly encode structured behavioral cues to achieve robust, human-like inference.

视觉推理社交线索多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。