arXiv:2504.05253cs.CV2025-04ICML被引 12

人类视觉依赖轮廓整合,大模型需海量数据才接近人眼表现

Contour Integration Underlies Human-Like Vision

  • 通过控制物体碎片化程度测试模型对轮廓的整合能力
  • 超1000个模型仅在约50亿数据训练下接近人类准确率
  • 具备方向性轮廓偏好且训练数据量大的模型表现更优

尽管深度学习在计算机视觉中取得巨大成功,模型在泛化到新输入分布时仍远逊于人类。现有基准未系统分析模型在多种可控条件下的失效点。本研究通过设计实验,系统检验了模型在不同物体碎片化水平下的目标识别能力,聚焦于人类视觉的核心特征——轮廓整合。50名人类受试者即使在仅有少量物体轮廓时仍保持高准确率。相比之下,绝大多数测试过的超1000个模型对增加轮廓的敏感性显著不足,多数表现仅略高于随机水平。只有当训练数据规模达到约50亿(∼5B)时,模型性能才开始接近人类。重要的是,人类表现出轮廓整合偏差——更倾向于识别由方向性片段构成的物体。我们发现,具备此特性的模型表现更好,且该偏差随训练数据规模增大而增强;训练模型显式学习轮廓整合可大幅提升形状偏差。综上,轮廓整合是支撑物体视觉的核心机制,可能通过大规模数据学习获得。

原文摘要 · Abstract (English)

Despite the tremendous success of deep learning in computer vision, models still fall behind humans in generalizing to new input distributions. Existing benchmarks do not investigate the specific failure points of models by analyzing performance under many controlled conditions. Our study systematically dissects where and why models struggle with contour integration -- a hallmark of human vision -- by designing an experiment that tests object recognition under various levels of object fragmentation. Humans (n=50) perform at high accuracy, even with few object contours present. This is in contrast to models which exhibit substantially lower sensitivity to increasing object contours, with most of the over 1,000 models we tested barely performing above chance. Only at very large scales ($\sim5B$ training dataset size) do models begin to approach human performance. Importantly, humans exhibit an integration bias -- a preference towards recognizing objects made up of directional fragments over directionless fragments. We find that not only do models that share this property perform better at our task, but that this bias also increases with model training dataset size, and training models to exhibit contour integration leads to high shape bias. Taken together, our results suggest that contour integration is a hallmark of object vision that underlies object recognition performance, and may be a mechanism learned from data at scale.

轮廓整合视觉认知大模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。