测试大模型能否像人一样主动观察,发现它们表现远不如人类。
An Exam for Active Observers

- 设计17个任务强制模型多次看图而非只看一眼
- 顶尖模型最高仅正确10.6%,人类平均96.1%
- 模型写代码也靠不住,缺乏自主修正视觉错误能力
人类视觉是闭环过程:目光不断根据中间假设调整,而非依赖单一快照。数十年心理学与认知科学表明,主动观察对多种任务至关重要。当前多模态大模型是否具备主动观察,现有视觉-语言评测无法回答。我们提出ActiveVision基准,包含17项跨3类任务,迫使模型进行重复视觉感知而非一次静态描述。前沿模型在该基准上表现崩溃:评估中最高性能的GPT-5.5(最高推理强度)仅解决10.6%的题目,11项任务得分为零;即便在多数推理与编码榜单领先的Claude Fable 5,也仅能解决3.5%,远低于三名人类参与者平均96.1%的准确率。此外,即使允许模型自动生成并运行视觉代码,其代码在真实图像上仍不可靠,而检测这些失败本身又需要主动感知——这正是模型所缺乏的能力。结果表明当前多模态大模型缺乏稳健的主动视觉观察,亟需能闭合感知-推理循环的架构与训练目标。
原文摘要 · Abstract (English)
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。