测试发现大模型看图推理靠背景知识,而非真懂图形关系。
Do Vision-Language Models Really Understand Visual Language?
- 构建多领域合成与真实图表测试集,评估模型识图与推理能力。
- 模型能准确识别实体,但理解关系的能力明显不足。
- 性能优势源于背景知识捷径,非真正掌握视觉语言逻辑。
视觉语言是通过符号、形状和空间布局传递信息的沟通系统,图表是其典型代表,用图像表达复杂概念及其关系。图表的符号性给构建具备理解能力的模型带来巨大挑战。近期研究认为大视觉语言模型(LVLMs)可处理涉及图表的复杂推理任务。本文通过开发全面的测试套件,评估LVLM在图表理解上的表现。测试集包含多种问题,聚焦概念实体及其关系,覆盖合成与真实图表,涵盖多个领域,以检验模型的识别与推理能力。评估结果表明,尽管模型能准确识别和推理实体,但对关系的理解能力显著受限。进一步分析显示,模型在图表理解上表现出色主要依赖其背景知识作为捷径来推断关系信息。因此我们得出结论:LVLM对图表理解的真实能力有限,其看似出色的推理表现是一种由背景知识等混淆因素引发的假象。
原文摘要 · Abstract (English)
Visual language is a system of communication that conveys information through symbols, shapes, and spatial arrangements. Diagrams are a typical example of a visual language depicting complex concepts and their relationships in the form of an image. The symbolic nature of diagrams presents significant challenges for building models capable of understanding them. Recent studies suggest that Large Vision-Language Models (LVLMs) can even tackle complex reasoning tasks involving diagrams. In this paper, we investigate this phenomenon by developing a comprehensive test suite to evaluate the diagram comprehension capability of LVLMs. Our test suite uses a variety of questions focused on concept entities and their relationships over a set of synthetic as well as real diagrams across domains to evaluate the recognition and reasoning abilities of models. Our evaluation of LVLMs shows that while they can accurately identify and reason about entities, their ability to understand relationships is notably limited. Further testing reveals that the decent performance on diagram understanding largely stems from leveraging their background knowledge as shortcuts to identify and reason about the relational information. Thus, we conclude that LVLMs have a limited capability for genuine diagram understanding, and their impressive performance in diagram reasoning is an illusion emanating from other confounding factors, such as the background knowledge in the models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。