arXiv:2602.20918cs.AIcs.CL2026-02被引 1

研究视觉上下文对句子可接受性判断的影响,发现人类不受影响,大模型则表现更差。

Predicting Sentence Acceptability Judgments in Multimodal Contexts

  • 对比人类与大模型在图文上下文中对句子可接受性的判断差异
  • 大模型在无视觉信息时预测准确率更高,且与自身概率相关性更强
  • 不同大模型表现各异,通义千问最接近人类模式

以往研究探讨了深度神经网络(尤其是变换器)在独立或文档上下文中预测人类句子可接受性判断的能力。本文考察视觉图像(即视觉上下文)对人类和大语言模型(LLMs)判断的影响。结果表明,与文本上下文不同,视觉图像对人类可接受性评分几乎没有影响。然而,大语言模型仍表现出先前在文档上下文中观察到的压缩效应。各类大模型均能以高精度预测人类判断,但总体而言,去除视觉上下文后性能略有提升。不同大模型的判断分布各异,其中Qwen的模式最接近人类,其他模型则偏离较大。大模型生成的可接受性判断与其归一化对数概率高度相关,但在存在视觉上下文时相关性下降,表明模型内部表示与输出预测之间的差距在多模态情境下增大。实验揭示了人类与大模型在多模态语境中处理句子的异同点。

原文摘要 · Abstract (English)

Previous work has examined the capacity of deep neural networks (DNNs), particularly transformers, to predict human sentence acceptability judgments, both independently of context, and in document contexts. We consider the effect of prior exposure to visual images (i.e., visual context) on these judgments for humans and large language models (LLMs). Our results suggest that, in contrast to textual context, visual images appear to have little if any impact on human acceptability ratings. However, LLMs display the compression effect seen in previous work on human judgments in document contexts. Different sorts of LLMs are able to predict human acceptability judgments to a high degree of accuracy, but in general, their performance is slightly better when visual contexts are removed. Moreover, the distribution of LLM judgments varies among models, with Qwen resembling human patterns, and others diverging from them. LLM-generated predictions on sentence acceptability are highly correlated with their normalised log probabilities in general. However, the correlations decrease when visual contexts are present, suggesting that a higher gap exists between the internal representations of LLMs and their generated predictions in the presence of visual contexts. Our experimental work suggests interesting points of similarity and of difference between human and LLM processing of sentences in multimodal contexts.

自然语言理解大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。