arXiv:2409.05862cs.CV2024-09NeurIPS被引 17

对比人类与视觉模型在3D形状推理上的表现,发现人类显著更优。

Evaluating Multiview Object Consistency in Humans and Image Models

  • 设计零样本视觉任务,通过多视角图像判断物体是否相同
  • 收集500人共3.5万次数据,发现人类性能远超所有模型
  • 揭示人类在难题上投入更多时间,模型与人类有差异但相关

我们引入一个基准测试,直接评估人类观察者与视觉模型在3D形状推断任务中的对齐程度。借鉴认知科学实验设计,要求参与者在无先验知识的情况下,基于一组图像判断其中是否包含相同或不同物体,尽管存在显著视角变化。使用涵盖常见物体(如椅子)和抽象形状(即程序生成的“无意义”物体)的多样化图像,构建了超过2000个独特图像集。向超过500名参与者发放任务,收集3.5万次行为数据,包括明确的选择行为、反应时间与注视轨迹等中间指标。随后评估常见视觉模型(如DINOv2、MAE、CLIP)的表现。结果显示,人类在所有任务中均显著优于所有模型。采用多尺度评估方法,识别出模型与人类之间的异同:虽然两者性能呈正相关,但人类在困难试次上分配了更多处理时间。所有图像、数据与代码均可通过项目页面获取。

原文摘要 · Abstract (English)

We introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cognitive sciences which requires zero-shot visual inferences about object shape: given a set of images, participants identify which contain the same/different objects, despite considerable viewpoint variation. We draw from a diverse range of images that include common objects (e.g., chairs) as well as abstract shapes (i.e., procedurally generated `nonsense' objects). After constructing over 2000 unique image sets, we administer these tasks to human participants, collecting 35K trials of behavioral data from over 500 participants. This includes explicit choice behaviors as well as intermediate measures, such as reaction time and gaze data. We then evaluate the performance of common vision models (e.g., DINOv2, MAE, CLIP). We find that humans outperform all models by a wide margin. Using a multi-scale evaluation approach, we identify underlying similarities and differences between models and humans: while human-model performance is correlated, humans allocate more time/processing on challenging trials. All images, data, and code can be accessed via our project page.

3D推理人类对比视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。