arXiv:2606.05409cs.CVcs.CL2026-06

对比人与模型学习新视觉概念的能力,发现模型易过度泛化。

Would you still call this Dax? Novel Visual References in VLMs and Humans

论文配图:Would you still call this Dax? Novel Visual References in VLMs and Humans
图 1 · 摘自论文原文
  • 构建全新视觉概念数据集NVRD,测试模型在上下文中学新概念能力
  • 模型面对与预训练矛盾的新概念时难以习得,且泛化远超人类
  • 适合研究认知机制或视觉语言模型泛化能力的学者参考

视觉语言模型(VLMs)如同人类学习者,常需面对新视觉概念,但其在接触后如何将新视觉参照映射到语言仍缺乏研究,尤其当这些参照与预训练知识相悖时。为此,我们提出新颖视觉参照数据集(NVRD):包含19,176张图像,覆盖90个视觉概念,每个概念有最多20种渐进扰动版本,以探测泛化能力。不同于以往对熟悉概念的视觉增强研究,NVRD完全由全新、开放性刺激构建,模拟人类真正首次接触新概念的情境。我们评估了3个开源和2个闭源模型,并与2,400条人类判断进行直接比较,发现:(i)当新概念与预训练知识冲突时,模型在上下文中难以习得;(ii)尽管模型与人类对视觉扰动的敏感度相关,但模型显著过泛化,将已学标签扩展至人类拒绝的刺激。我们贡献NVRD作为人类与机器视觉概念学习研究的语料库与基准。

原文摘要 · Abstract (English)

Vision-language models (VLMs), like human learners, are frequently exposed to new visual concepts, but how they map novel visual references to language after exposure remains largely underexplored, particularly when those references contradict prior knowledge from pre-training. To study this, we present the Novel Visual References Dataset (NVRD): 19,176 images spanning 90 visual concepts across different levels of visual novelty, each with up to 20 increasingly perturbed versions of the original object to probe generalization. Unlike prior work on visual augmentations of familiar concepts, NVRD comprises entirely novel, open-ended stimuli constructed from scratch, mirroring how humans encounter genuinely new concepts. We evaluate 3 open- and 2 closed-source models alongside 2,400 human judgments for direct human-model comparison, and find that (i) models struggle to acquire novel concepts in-context when they contradict prior knowledge, and (ii) while models and humans show correlated sensitivity to visual perturbations, models significantly overgeneralize, extending learned labels to stimuli that humans reject. We contribute NVRD as a corpus and benchmark for research on visual concept learning in both humans and machines.

视觉语言模型概念学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。