arXiv:2601.18065cs.CL2026-01ACL

视觉语言模型更像人类,对词语具体性的敏感度更高。

Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision-Language Models

  • 用相同文本骨干对比图文模型与纯文本模型。
  • 图文模型在具体词汇上表现更好,且表征结构更符合具体性规律。
  • 适合研究多模态学习与人类认知对齐的学者。

视觉语言模型(VLMs)在仅用文本提示评估时,是否比纯文本大语言模型(LLMs)展现出更接近人类的词汇具体性敏感性?我们通过在多个模型规模下,对比相同Llama文本骨干与其对应的Llama Vision版本,将多模态预训练视为感知基础化的消融实验而非推理时访问图像。从三个互补层面测量具体性效应:(i) 输出行为,考察问题具体性与问答准确率的关系;(ii) 嵌入几何,检验表征是否沿具体性轴组织;(iii) 注意力动态,通过注意力熵衡量上下文依赖程度。此外,我们提取模型的词级具体性评分,评估其与人类标准分布的一致性,测试多模态训练是否带来更贴近人类的判断。在多个基准和规模下,VLMs 在具体输入上表现出更大提升,具有更清晰的具体性结构化表征,评分更接近人类分布,注意力模式也显示更强的感知根基特征。

原文摘要 · Abstract (English)

Do vision--language models (VLMs) develop more human-like sensitivity to linguistic concreteness than text-only large language models (LLMs) when both are evaluated with text-only prompts? We study this question with a controlled comparison between matched Llama text backbones and their Llama Vision counterparts across multiple model scales, treating multimodal pretraining as an ablation on perceptual grounding rather than access to images at inference. We measure concreteness effects at three complementary levels: (i) output behavior, by relating question-level concreteness to QA accuracy; (ii) embedding geometry, by testing whether representations organize along a concreteness axis; and (iii) attention dynamics, by quantifying context reliance via attention-entropy measures. In addition, we elicit token-level concreteness ratings from models and evaluate alignment to human norm distributions, testing whether multimodal training yields more human-consistent judgments. Across benchmarks and scales, VLMs show larger gains on more concrete inputs, exhibit clearer concreteness-structured representations, produce ratings that better match human norms, and display systematically different attention patterns consistent with increased grounding.

视觉语言模型具体性认知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。