arXiv:2606.03142cs.CV2026-06被引 1

区分视觉与事实正确性,揭示大模型看图能力的真相

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy

论文配图:Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy
图 1 · 摘自论文原文
  • 设计新评估框架,分离视觉理解与记忆事实的影响
  • 15个主流模型中多数依赖事实而非真实看图,部分表现受提示误导
  • 提出可量化模型倾向性的指标,适合评估图表分析可靠性

大型视觉-语言模型(LVLMs)具备强大的图像解读能力,但其回答是基于对视觉证据的真实推理,还是源于训练中习得的事实先验尚不明确。现有评估混淆了两种来源,掩盖了正确视觉解释被记忆事实覆盖的情况。本文提出一个框架,将视觉正确性与事实正确性分离,揭示现有可视化素养评估的有效性局限。在三个实验中测试15个前沿LVLMs:(1) 多数模型在标准测试(VLAT)上达到人类水平,但这可能反映的是事实回忆而非视觉理解;而在随机数据测试(reVLAT)中,当视觉理解被事实先验压制时,其表现被低估。(2) 使用反事实可视化素养评估测试(CVLAT)及能力归一化仲裁指标,按视觉-事实依赖指数(VFRI)分类模型,发现以视觉为导向的多数和以事实为导向的少数,少数接近零值需警惕;30人基准测试显示,人类在冲突时仍优先遵循图表,提供人类参照。(3) 提示干预可改变优先级,但效果高度依赖模型且方向不对称,高读图能力不预示提示可控性。总体而言,高可视化准确率不足以证明忠实的视觉推理:将模型可靠集成于可视化分析,需同时评估其在视觉与事实冲突时的权衡能力。基准与代码:https://github.com/JaeyoungKim-HCIL/CVLAT

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) show strong visualization interpretation, yet it is unclear whether their responses reflect genuine reasoning over visual evidence or factual priors learned during training. Current evaluations mix these two sources, obscuring when correct visual interpretation is overridden by memorized facts. We present a framework that isolates visual correctness from factual correctness, revealing validity limitations in existing visualization literacy assessments. Across three experiments with 15 state-of-the-art LVLMs: (1) several models reach human-level performance on standard tests (VLAT), but this may reflect factual recall rather than visual understanding, while randomized-data tests (reVLAT) underestimate literacy when correct visual interpretation is superseded by factual priors. (2) Using our Counterfactual Visualization Literacy Assessment Test (CVLAT) with capability-normalized arbitration metrics, we classify models by the sign of their visual-factual reliance index (VFRI), revealing a visualization-oriented majority and a factual knowledge-oriented minority, though several near-zero cases warrant caution. A human baseline (N=30) on the same counterfactual items confirms that people overwhelmingly follow the chart under conflict, providing a human reference point. (3) Prompt-based intervention can shift prioritization, but its effectiveness is highly model-dependent and direction-asymmetric, and high chart-reading capability does not predict prompt-controllability. Overall, high visualization accuracy is not sufficient evidence of faithful visual reasoning: reliable integration into visual analytics requires evaluating not only visualization literacy but also how models arbitrate between visual evidence and factual priors when the two diverge. Benchmark and code: https://github.com/JaeyoungKim-HCIL/CVLAT

视觉理解模型评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。