arXiv:2606.19552cs.CL2026-06

构建视觉语言结构歧义评测集,检验模型看图解歧能力

LaViSA: A Language and Vision Structural Ambiguity Benchmark

论文配图:LaViSA: A Language and Vision Structural Ambiguity Benchmark
图 1 · 摘自论文原文
  • 设计包含七类歧义的图文匹配数据集,评估模型结合图像理解句子歧义
  • 多模型测试显示当前大模型仍难处理部分复杂歧义和细微语义差异
  • 适合研究视觉语言理解、模型可解释性与多模态推理的研究者使用

结构歧义指单个句子因句法结构导致多种合理解释,是语言理解的根本挑战。视觉场景能为消歧提供有效线索,而视觉语言模型需具备从图像中推导可能语义解释的能力。本文提出语言与视觉结构歧义基准(LaViSA),用于评估模型利用视觉场景解决结构歧义的能力。LaViSA包含七类歧义句、对应的消歧句及其对应图像。通过该基准对多种视觉语言模型(含闭源与开源、不同参数量与推理能力)进行系统评估。实验表明,尽管近期模型能部分借助视觉信息消歧,但在某些歧义类型及视觉上细微的语义差异上仍表现不佳,揭示了当前模型在利用视觉场景解决结构歧义方面仍存在明显局限。

原文摘要 · Abstract (English)

Structural ambiguity arises when a single sentence admits multiple valid interpretations due to its syntactic structure, posing a fundamental challenge for language understanding. Visual scenes serve as useful cues for resolving such ambiguity, and Vision and Language Models (VLMs) need to be capable of deriving possible semantic interpretations from visual scenes. We introduce Language and Vision Structural Ambiguity (LaViSA), a benchmark designed to evaluate the ability of VLMs to resolve structural ambiguity leveraging visual scenes. LaViSA consists of ambiguous sentences, their disambiguated sentences, and corresponding images of these disambiguated sentences across seven ambiguity categories. Using LaViSA, we conduct a comprehensive evaluation of diverse VLMs, including both proprietary and open-source models with varying parameter scales and reasoning capabilities. Experimental results show that although recent VLMs can leverage visual scenes to resolve structural ambiguity to a some extent, they still struggle with certain ambiguity types and visually subtle semantic distinctions, indicating remaining limitations in resolving structural ambiguity using visual scenes.

视觉语言结构歧义多模态评测模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。