视觉语言模型难以理解场景的使用功能,因缺乏身体经验。
The Limits of Learning from Pictures and Text: Vision-Language Models and Embodied Scene Understanding
- 用人类响应分布校准的相似度指标评估模型表现
- 模型在物体功能理解上明显落后于人类,且新模型未改善
- 图像描述数据中缺少对使用者的描述,导致知识缺失
什么信息足以习得人类场景理解的全部丰富性?分布假设认为语言与图像的统计共现能捕捉视觉认知背后的概念知识。视觉语言模型(VLMs)在大规模图文配对数据上训练,但缺乏具身经验,是检验该假设的理想对象。我们通过两个实验,将18个VLM生成的描述与超过2000名人类观察者在15个高层次场景理解任务中的回答进行比较,涵盖常识、使用功能、感官体验、情感反应和未来预测。由于许多任务无标准答案,我们提出了人类校准余弦距离(HCD)指标,衡量模型输出与人类响应分布的相似性,并以人内变异性为尺度。实验1显示,VLM在常识任务接近人类水平,但在使用功能任务存在显著缺陷,且不受提示工程影响,新版本模型亦未改进。实验2测试了六种机制假说,发现该缺陷是结构性而非风格性,提供显式空间信息也无法解决。语料库分析表明,图像描述数据中关于使用功能的语言极少提及主体,符合格赖斯语用学解释——具身知识在语言中系统性缺失。综合表明,仅靠图像与文本的分布学习不足以实现基于功能的场景理解,部分人类视觉认知维度可能需要照片或描述无法编码的主体中心、三维体感经验。
原文摘要 · Abstract (English)
What information is sufficient to learn the full richness of human scene understanding? The distributional hypothesis holds that the statistical co-occurrence of language and images captures the conceptual knowledge underlying visual cognition. Vision-language models (VLMs) are trained on massive paired text-image corpora but lack embodied experience, making them an ideal test of the distributional hypothesis. We report two experiments comparing descriptions generated by 18 VLMs to those of over 2000 human observers across 15 high-level scene understanding tasks, spanning general knowledge, affordances, sensory experiences, affective responses, and future prediction. Because many tasks lack ground truth answers, we developed a Human-Calibrated Cosine Distance (HCD) metric that measures VLM output similarity to the distribution of human responses, scaled by within-human variability. In Experiment 1, VLMs approached human-level performance on general knowledge tasks, but showed a robust deficit for affordance tasks that resisted prompt engineering and did not improve with newer model releases. In Experiment 2, we tested six mechanistic hypotheses for explaining this affordance gap, finding that the deficit was structural rather than stylistic and was not resolved by providing explicit spatial information. Corpus analyses revealed that image captioning datasets contain sparse agent-addressed affordance language, consistent with Gricean accounts of why embodied knowledge may be systematically underrepresented in language. Together, these findings suggest that distributional learning from images and text is insufficient for affordance-based scene understanding, implying that some dimensions of human visual cognition may require the kind of agent-centered, three-dimensional experience that no photograph or caption can encode.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。