测试大模型是看图推理还是靠记忆答的,发现多数依赖记忆而非真实视觉理解。
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
- 用AI自动修改图像细节,控制性地改变视觉线索。
- 在755个测试样本上,模型准确率下降超40%,说明依赖记忆而非看图。
- 首个评估多模态模型视觉可信度的基准,适合研究模型可解释性的人。
近期研究显示,引入长链思维(CoT)可显著提升多模态大模型(MLLMs)解决复杂问题的能力,但其有效性机制尚不明确。本文提出一种基于GPT-Image-1的提示驱动、可控制的图像编辑流水线,实现对特定视觉线索的精准修改。在此基础上构建了首个评估MLLMs视觉推理能力与视觉忠实度的基准VFaith-Bench,包含755个样本,分为五个子集及一个人工标注的感知任务。通过修改关键视觉特征并生成对比问答对,测试模型在不同图像细节下的表现。结果表明,主流模型在图像修改后平均准确率下降超过40%,揭示其推理主要依赖记忆而非真实视觉感知。同时设计量化指标分析推理来源,深入验证多个主流模型系列与开源推理模型的表现差异。
原文摘要 · Abstract (English)
Recent extensive works have demonstrated that by introducing long CoT, the capabilities of MLLMs to solve complex problems can be effectively enhanced. However, the reasons for the effectiveness of such paradigms remain unclear. It is challenging to analysis with quantitative results how much the model's specific extraction of visual cues and its subsequent so-called reasoning during inference process contribute to the performance improvements. Therefore, evaluating the faithfulness of MLLMs' reasoning to visual information is crucial. To address this issue, we first present a cue-driven automatic and controllable editing pipeline with the help of GPT-Image-1. It enables the automatic and precise editing of specific visual cues based on the instruction. Furthermore, we introduce VFaith-Bench, the first benchmark to evaluate MLLMs' visual reasoning capabilities and analyze the source of such capabilities with an emphasis on the visual faithfulness. Using the designed pipeline, we constructed comparative question-answer pairs by altering the visual cues in images that are crucial for solving the original reasoning problem, thereby changing the question's answer. By testing similar questions with images that have different details, the average accuracy reflects the model's visual reasoning ability, while the difference in accuracy before and after editing the test set images effectively reveals the relationship between the model's reasoning ability and visual perception. We further designed specific metrics to expose this relationship. VFaith-Bench includes 755 entries divided into five distinct subsets, along with an additional human-labeled perception task. We conducted in-depth testing and analysis of existing mainstream flagship models and prominent open-source model series/reasoning models on VFaith-Bench, further investigating the underlying factors of their reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。