arXiv:2606.25445cs.CVcs.AI2026-06

首个面向真实场景变化的图文描述评估基准,揭示现有模型在复杂情境下的系统性失效。

C3-Bench: A Context-Aware Change Captioning Benchmark

论文配图:C3-Bench: A Context-Aware Change Captioning Benchmark
图 1 · 摘自论文原文
  • 构建包含4996对图像与51类真实变化场景的多领域评测集
  • 提出基于大模型评分的细粒度评估框架,首次引入可逆性指标衡量理解一致性
  • 发现主流模型在非训练场景下严重退化,适合关注模型鲁棒性的研究者

尽管变化描述系统受到广泛关注以应对不断变化的世界,但其在多样真实变化情境中的实际表现仍缺乏全面评估框架。为此,我们提出C3-Bench,一个面向上下文感知变化描述的综合性评测基准。C3-Bench包含:(1)4,996对人类标注的图像对,涵盖自然场景、遥感影像、图像编辑和异常等四个领域的51种真实变化情境,每类均来自多个以变化为核心的社区,经过精心设计;(2)首个在变化描述任务中应用大模型作为评判者的评估框架,可衡量正确性、具体性、流畅性和相关性等细粒度维度,并引入新颖的可逆性指标,探索模型是否具备对称一致的变化理解能力。基于C3-Bench,我们对32个模型进行了评测——包括传统变化描述模型、专有大型多模态模型(LMMs)以及2B至90B规模的开源LMMs。结果揭示了当前变化描述范式的根本盲区:一旦变化情境偏离训练设定,传统模型性能急剧下降,甚至最先进的LMM如GPT-5.2也表现出显著的领域与位置依赖性错误,导致可靠变化理解被扭曲。通过显式揭示这些隐藏的失败模式并实现可测量化,我们明确了构建通用且可信变化描述系统的新前沿。所有代码与数据集均公开于项目页面。

原文摘要 · Abstract (English)

While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contexts remains largely unexplored due to the lack of comprehensive evaluation frameworks. To fill this gap, we propose C3-Bench, a comprehensive benchmark for evaluating Context-aware Change Captioning. C3-Bench features: (1) 4,996 human-labeled image pairs of 51 real-world change contexts across four domains (e.g., natural scenes, remote sensing imagery, image editing, and anomalies), each with diverse, carefully curated scenarios derived from multiple change-centric communities; and (2) the first LLM-as-Judge evaluation framework in the change captioning task that measure fine-grained dimensions (e.g., correctness, specificity, fluency, and relevance), along with a novel reversibility metric exploring whether models understand changes with symmetric consistency. Based on C3-Bench, we benchmark 32 models -- including conventional change captioning models, proprietary Large Multimodal Models (LMMs), and 2B-90B open-source LMMs. We reveal a fundamental blind spot in the prevailing change captioning paradigm: Once the change context departs from training-style regimes, conventional models collapse, and even state-of-the-art LMMs such as GPT-5.2 exhibit systematic domain- and position-dependent errors that distort reliable change understanding. By making these hidden failure modes explicit and measurable, we delineate the next frontier for building generalizable and trustworthy change captioning systems. All codes and datasets are publicly available on the project page.

变化描述评测基准大模型评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。