测试视觉语言模型能否推断说话人无知,发现仅Claude展现多线索融合能力。
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues
- 通过视觉与语言线索组合,测试模型对说话人无知的推断能力
- 仅提供视觉线索时,三模型均依赖字面意义解释数字
- Claude是唯一能融合双线索的模型,显示初步语用推理潜力
本研究探究视觉语言模型(VLMs)是否具备语用推理能力,聚焦于暗示说话人缺乏精确知识的隐含含义。通过系统操控上下文线索:视觉呈现的情境(视觉线索)和基于问题焦点(QUD)的语言提示(语言线索)。当仅提供视觉线索时,三种先进VLM(GPT-4o、Gemini 1.5 Pro、Claude 3.5 sonnet)的解释主要依据被修改数字的词汇意义。加入语言线索以增强上下文信息后,Claude展现出更接近人类的推理,能整合两类线索。而GPT与Gemini仍偏好精确、字面的解释。尽管上下文线索影响增强,它们仍独立处理各线索,并与语义特征对齐,而非进行上下文驱动的推理。结果表明,尽管模型对线索处理方式不同,但Claude融合多线索的能力可能预示多模态模型中语用能力的萌芽。
原文摘要 · Abstract (English)
This study investigates whether vision-language models (VLMs) can perform pragmatic inference, focusing on ignorance implicatures, utterances that imply the speaker's lack of precise knowledge. To test this, we systematically manipulated contextual cues: the visually depicted situation (visual cue) and QUD-based linguistic prompts (linguistic cue). When only visual cues were provided, three state-of-the-art VLMs (GPT-4o, Gemini 1.5 Pro, and Claude 3.5 sonnet) produced interpretations largely based on the lexical meaning of the modified numerals. When linguistic cues were added to enhance contextual informativeness, Claude exhibited more human-like inference by integrating both types of contextual cues. In contrast, GPT and Gemini favored precise, literal interpretations. Although the influence of contextual cues increased, they treated each contextual cue independently and aligned them with semantic features rather than engaging in context-driven reasoning. These findings suggest that although the models differ in how they handle contextual cues, Claude's ability to combine multiple cues may signal emerging pragmatic competence in multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。