构建首个细粒度图像差异描述综合基准,提升模型对细微变化的感知与表达能力。
OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference Captioning
- 提出多尺度差异感知模块,增强模型定位和描述图像间微小变化的能力。
- 在324个复杂场景下,平均60词标注覆盖12类变化,数据更全面。
- 适合图像差异分析、跨场景视觉理解等方向的研究者使用。
图像差异描述(IDC)旨在生成图像对之间细微差异的自然语言描述,需精准定位视觉变化并保持语义连贯。现有数据集普遍存在广度与深度不足的问题:(1)样本局限于特定场景中的有限对象变化;(2)描述过于简单。为此,我们构建了OmniDiff,一个涵盖324种多样化场景(含真实世界与3D合成环境)的综合性数据集,提供平均60词长度的细粒度人工标注,覆盖12种不同变化类型。基于此,我们提出M$^3$Diff,一种融合可插拔多尺度差异感知(MDP)模块的多模态大模型,显著提升差异识别与描述准确性,同时保持模型泛化能力。在Spot-the-Diff、IEdit、CLEVR-Change、CLEVR-DC及OmniDiff等多个基准上,M$^3$Diff达到最新性能,跨场景差异识别准确率显著优于现有方法。数据集、代码与模型将公开发布,以推动该领域研究。
原文摘要 · Abstract (English)
Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements, existing datasets often lack breadth and depth, limiting their applicability in complex and dynamic environments: (1) from a breadth perspective, current datasets are constrained to limited variations of objects in specific scenes, and (2) from a depth perspective, prior benchmarks often provide overly simplistic descriptions. To address these challenges, we introduce OmniDiff, a comprehensive dataset comprising 324 diverse scenarios-spanning real-world complex environments and 3D synthetic settings-with fine-grained human annotations averaging 60 words in length and covering 12 distinct change types. Building on this foundation, we propose M$^3$Diff, a MultiModal large language model enhanced by a plug-and-play Multi-scale Differential Perception (MDP) module. This module improves the model's ability to accurately identify and describe inter-image differences while maintaining the foundational model's generalization capabilities. With the addition of the OmniDiff dataset, M$^3$Diff achieves state-of-the-art performance across multiple benchmarks, including Spot-the-Diff, IEdit, CLEVR-Change, CLEVR-DC, and OmniDiff, demonstrating significant improvements in cross-scenario difference recognition accuracy compared to existing methods. The dataset, code, and models will be made publicly available to support further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。