arXiv:2606.07872cs.CV2026-06被引 1

测试多模态模型是否真依赖关键视觉证据做推理

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?

论文配图:VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?
图 1 · 摘自论文原文
  • 设计配对图像基准,微调视觉证据使答案必然翻转
  • 24个模型中多数在证据变化时仍重复相同答案
  • 揭示准确率不等于视觉接地,适合评估模型可靠性

当多模态大语言模型正确回答视觉推理问题时,其预测是否真正基于任务关键的视觉证据?正确答案可能伴随错误推理,仅靠准确率无法全面检验模型的视觉接地能力。我们提出VisualFLIP,一个包含1,374张图像的成对基准,覆盖基数、属性、空间和逻辑四类任务。每对图像保持问题一致,仅微调视觉证据,使正确答案必然翻转。我们评估24个多模态大模型(MLLMs),使用配对准确率(需同时正确解答两图)与崩溃率(CR,即至少解对一图但两图答案相同且非空的比例)。结果表明,配对正确性与证据依赖性相关但独立:部分能力强的模型在关键视觉变化后仍无法更新输出;在序列设置下,若修改后的图像紧随先前答案,崩溃现象更严重。详细信息见项目页:https://didizhu-judy.github.io/VisualFLIP/

原文摘要 · Abstract (English)

When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence? Correct answers can coexist with flawed reasoning, making accuracy alone an incomplete test of grounding. We introduce VisualFLIP, a paired benchmark with 1,374 images arranged as same-question perturbation pairs across cardinality, attribute, spatial, and logic tasks. Each pair keeps the question fixed but minimally changes the evidence so the gold answer deterministically flips. We evaluate 24 MLLMs with pair accuracy, which requires solving both sides of a pair, and Collapse Rate (CR), which measures how often a model that solves at least one side repeats the same non-empty answer for both images. Together, these metrics show that paired correctness and evidence dependence are related but distinct: capable models can still fail to update after task-critical visual changes, and collapse becomes more severe for some models when the edited image follows an earlier answer in a sequential setting. Further details are available on our project page: https://didizhu-judy.github.io/VisualFLIP/

多模态视觉推理模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。