评测大模型对图像细微差异的识别能力,发现其在噪声纹理等低级变化上表现很差。
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification

- 构建包含1756个四选一题目的细粒度图像差异识别基准
- 小模型在噪声纹理等低级变化上准确率仅8.7%-33.3%
- 揭示主流模型在对比视觉理解上的盲区,适合评估模型诊断
多模态大语言模型在通用视觉理解任务中表现优异,但在识别两幅相似图像间的细微差异这一基础能力上仍显不足。本文提出VDiff-Bench,一个面向细粒度图像差异识别的挑战性多选基准,包含1,756个四选一问题,覆盖10类变化:位置、运动、局部/整体颜色、外观/消失、噪声/分辨率、纹理、替换/大小、OCR/文本、光照。每道题对应一对图像及四个选项:真实差异、两个难负例描述和一个“无差异”干扰项。为增加难度,负例经过精心设计,需模型区分实际变化与语义相近的干扰项。对11个先进开源与闭源MLLM的实验表明,细粒度视觉比较仍极脆弱:模型在不同来源和变化类别间表现不均,对细微低级变化如噪声和纹理持续失败。例如,三个7-8B规模的开源模型在语义变化上得分52.5%-70.6%,但在噪声与纹理变化上仅8.7%-33.3%,常错误判断为无变化。令人意外的是,尽管其他闭源模型表现良好,Grok 4.3在噪声与纹理差异识别上出现显著性能下降,远低于Kimi K2.5和K3等大型开源模型。VDiff-Bench为诊断MLLM的对比视觉理解能力提供了精准工具,暴露出单图像任务无法捕捉的缺陷。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed between two similar images. We introduce VDiff-Bench, a challenging multiple-choice benchmark for fine-grained Image Difference Identification. VDiff-Bench contains 1,756 four-way questions over image pairs and covers 10 change categories: position, motion, regional image color, overall image color, appearance/disappearance, noise/resolution, texture, substitution/size, OCR/text, and illumination. Each question corresponds to two image inputs with 4 choices: the true difference, two hard negative descriptions, and a "no difference" distractor. To make the task challenging, we specifically curate ground-truth-conditioned negatives that require models to distinguish the actual change from nearby semantic alternatives. Experiments with 11 state-of-the-art open- and closed-source MLLMs show that fine-grained visual comparison remains brittle: models exhibit uneven performance across sources and change categories, with persistent failures on subtle low-level changes like noises and textures. For instance, three 7-8B-scale open-source MLLMs score 52.5-70.6% on semantic changes but only 8.7-33.3% on low-level changes like noise and texture, falsely assuming no changes between two image inputs. Surprisingly, despite strong performance of other closed-source commercial models, Grok 4.3 demonstrate remarkable performance drop on identifying noise and texture differences between images, falling significantly behind large open-source models like Kimi K2.5 and K3. Overall, VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。