arXiv:2608.10195cs.CVcs.LG2026-08中稿 · VISxVision 2026, a…

测试视觉模型是否像人一样分组视觉信息,发现传统指标会高估其感知能力。

More Accurate, Less Human: Gestalt Grouping in Vision Models

论文配图:More Accurate, Less Human: Gestalt Grouping in Vision Models
图 1 · 摘自论文原文
  • 设计四类分组任务,对比模型与人类的感知一致性。
  • 45个模型中多个闭源模型与人类响应差异显著,但准确率仍很高。
  • 为可视化研究提供无需用户实验的可复用评估工具。

人类视觉会将看到的内容组织成整体:相同颜色的点形成序列,相似标记聚成类别,形状构成可识别物体。这些是可视化设计所依赖的格式塔规律。现有视觉模型是否具备此类组织能力尚未系统检验。本文引入一套行为测试集,基于已有感知研究的人类数据,评估模型在四类分组任务中的表现:标记-颜色异常检测、颜色序列计数、轮廓识别和物体异常检测。测试覆盖45个模型,包括监督、自监督、对比学习视觉语言编码器、开源大模型及闭源基础模型。结果表明,模型与人类响应的一致性揭示了传统性能指标无法区分的感知组织特征;多个闭源模型虽有较高基准准确率,但与人类的对齐程度明显偏低。该评估方法为可视化研究提供了一个可复用的基准,无需开展新用户实验,即可审计进入可视化流程的模型是否以人类观众的方式理解视觉内容。

原文摘要 · Abstract (English)

Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.

视觉感知模型评估格式塔

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。