arXiv:2504.16060cs.CL2025-04被引 7

现有视觉语言模型在指代表达生成中缺乏实际沟通能力。

Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation

  • 从合作沟通原则出发,构建新数据集评估模型指代能力
  • 发现模型存在指代不唯一、信息冗余、空间提示不足三大问题
  • 现有自动评测无法捕捉这些缺陷,需改进评价体系

指代表达生成(REG)是评估视觉语言模型实用沟通能力的核心任务,不仅要求语义准确,还需遵循合作沟通原则(Grice, 1975)。然而当前对视觉语言模型(VLMs)的评估常忽略语用维度,将REG简化为区域描述任务,忽视格赖斯准则。本文从语用角度重新审视REG,构建包含1500张图像的新数据集RefOI,涵盖书面与口语指代表达。系统评估前沿VLMs后发现其在语用能力上存在三类关键缺陷:(1)无法唯一确定所指对象;(2)包含过多或无关信息;(3)与人类语用偏好不符,如空间线索使用不足。同时表明标准自动评估无法识别这些语用错误,反而强化表面线索。研究呼吁建立更符合人类真实交流的模型与评估框架。

原文摘要 · Abstract (English)

Referring Expression Generation (REG) is a core task for evaluating the pragmatic competence of vision-language systems, requiring not only accurate semantic grounding but also adherence to principles of cooperative communication (Grice, 1975). However, current evaluations of vision-language models (VLMs) often overlook the pragmatic dimension, reducing REG to a region-based captioning task and neglecting Gricean maxims. In this work, we revisit REG from a pragmatic perspective, introducing a new dataset (RefOI) of 1.5k images annotated with both written and spoken referring expressions. Through a systematic evaluation of state-of-the-art VLMs, we identify three key failures of pragmatic competence: (1) failure to uniquely identify the referent, (2) inclusion of excessive or irrelevant information, and (3) misalignment with human pragmatic preference, such as the underuse of minimal spatial cues. We also show that standard automatic evaluations fail to capture these pragmatic violations, reinforcing superficial cues rather than genuine referential success. Our findings call for a renewed focus on pragmatically informed models and evaluation frameworks that align with real human communication.

指代生成语用能力视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。