arXiv:2504.09809cs.HCcs.AI2025-04被引 10

发现大模型答可视化题靠记忆而非看图,提出新评估框架验证真看懂了没。

See or Recall: A Sanity Check for the Role of Vision in Solving Visualization Question Answer Tasks with Multimodal LLMs

  • 用规则树和检查表分离视觉感知与知识回忆的影响
  • 无图像时模型仍能正确回答超半数题目,依赖的是记忆而非看图
  • 适合研究多模态模型评估、可视化理解或避免误判的学者

多模态大语言模型(MLLM)已具备联合处理视觉与语言的能力,可对各类数据可视化进行理解和问答。当前主流评估方式是衡量模型在可视化问答(VisQA)任务中的推理能力,类比人类可视化素养。然而我们发现,模型在面对可视化问题时的推理机制与人类存在根本差异:即使不提供图像,模型仍能正确回答大量问题,无论是否有选项。这表明模型可能主要依赖其内部存储的海量知识进行事实回忆,而非依赖视觉信号。这一现象揭示现有评估可能无法真实反映模型的视觉推理能力。为此,我们提出一个综合性的合理性检验框架,结合规则决策树与合理性检查表,以分离‘看见’(视觉处理)与‘回忆’(先验知识依赖)的作用。该框架可验证现有VisQA数据集的有效性,识别模型是真正通过视觉理解、受记忆干扰,还是依赖归纳偏见作答。本研究强调,在使用MLLM进行可视化理解研究时,必须审慎设计评估方法。

原文摘要 · Abstract (English)

Recent developments in multimodal large language models (MLLM) have equipped language models to reason about vision and language jointly. This permits MLLMs to both perceive and answer questions about data visualization across a variety of designs and tasks. Applying MLLMs to a broad range of visualization tasks requires us to properly evaluate their capabilities, and the most common way to conduct evaluation is through measuring a model's visualization reasoning capability, analogous to how we would evaluate human understanding of visualizations (e.g., visualization literacy). However, we found that in the context of visualization question answering (VisQA), how an MLLM perceives and reasons about visualizations can be fundamentally different from how humans approach the same problem. During the evaluation, even without visualization, the model could correctly answer a substantial portion of the visualization test questions, regardless of whether any selection options were provided. We hypothesize that the vast amount of knowledge encoded in the language model permits factual recall that supersedes the need to seek information from the visual signal. It raises concerns that the current VisQA evaluation may not fully capture the models' visualization reasoning capabilities. To address this, we propose a comprehensive sanity check framework that integrates a rule-based decision tree and a sanity check table to disentangle the effects of "seeing" (visual processing) and "recall" (reliance on prior knowledge). This validates VisQA datasets for evaluation, highlighting where models are truly "seeing", positively or negatively affected by the factual recall, or relying on inductive biases for question answering. Our study underscores the need for careful consideration in designing future visualization understanding studies when utilizing MLLMs.

多模态模型可视化问答评估基准认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。