arXiv:2605.04503cs.CVcs.AI2026-05

构建首个全面且严谨的图像差异描述评测基准,解决现有方法评估不准的问题。

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

论文配图:DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
图 1 · 摘自论文原文
  • 设计涵盖10类差异的复杂评测集,提升任务多样性与挑战性。
  • 用大模型作为裁判,结合人工验证差异列表,实现更可靠的评估。
  • 发现开源模型在推理能力上显著落后,适合研究视觉-语言模型性能边界。

图像差异描述(IDC)旨在生成精确描述两张图像间差异的自然语言文本,是细粒度变化感知、跨模态推理和图像编辑数据构建的关键评测任务。然而,现有基准缺乏多样性与组合复杂性,且常用词汇重叠指标(如BLEU、METEOR)无法捕捉语义一致性或惩罚幻觉,导致对多模态大模型(MLLMs)在IDC任务上的评估不全面、不可靠。为此,我们提出DiffCap-Bench,一个覆盖十类不同差异类型的综合性评测基准,确保任务多样性和组合复杂性。同时,我们设计基于人工验证差异列表的LLM-as-a-Judge评估协议,实现对模型捕获与描述视觉变化能力的稳健评估。通过对顶尖MLLMs的广泛测试,我们揭示了专有模型与开源模型间的显著性能差距,强调了推理能力的关键作用,并识别出模型扩展中的明确局限。该框架与人类专家判断高度一致,且与下游图像编辑数据质量有强相关性。这些发现确立了DiffCap-Bench作为可靠评测框架及下游应用预测工具的价值。基准与代码将公开可用,以支持后续研究。

原文摘要 · Abstract (English)

Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change perception, cross-modal reasoning, and image editing data construction. However, existing benchmarks lack diversity and compositional complexity, and standard lexical-overlap metrics (e.g., BLEU, METEOR) fail to capture semantic consistency or penalize hallucinations, which together prevent a comprehensive and robust evaluation of multimodal large language models (MLLMs) on IDC. To address these gaps, we introduce DiffCap-Bench, a comprehensive IDC benchmark covering ten distinct difference categories to ensure diversity and compositional complexity. Furthermore, we propose an LLM-as-a-Judge evaluation protocol grounded in human-validated Difference Lists, enabling a robust assessment of models' ability to both capture and describe visual changes. Through extensive evaluation of state-of-the-art MLLMs, we reveal significant performance gaps between proprietary and open-source models, highlight the critical importance of reasoning capability, and identify clear limitations in model scaling. Our framework also demonstrates strong alignment with human expert judgments and strong correlation with downstream image editing data construction quality. These findings establish DiffCap-Bench as both a reliable IDC evaluation framework and a practical predictor of downstream utility. The benchmark and code will be made publicly available to support further research.

图像理解多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。