arXiv:2412.15939cs.CVcs.AI2024-12中稿 · the IEEE/CVF Winte…被引 6

用合成数据提升真实图像差异描述能力,效果显著优于传统方法。

Reframing Image Difference Captioning with BLIP2IDC and Synthetic Augmentation

  • 将BLIP2改造为BLIP2IDC,低成本适配图像差异描述任务
  • 合成数据增强使模型在真实世界数据集上性能大幅提升
  • 提出新数据集Syned1,适合训练和评估高阶差异描述模型

近年来生成模型的发展使得大规模图像编辑变异成为可能。为应对该技术带来的潜在危害,图像差异描述(IDC)任务旨在描述两张图像之间的差异。尽管该任务在简单3D渲染图像上表现良好,但在真实世界图像上仍面临挑战,主要源于训练数据稀缺和捕捉复杂图像细粒度差异的困难。为此,本文提出一种简单而有效的框架,既能将现有图像字幕模型适配至IDC任务,又能扩充数据集。我们引入了低成本适配的BLIP2IDC,其在真实世界IDC数据集上的表现显著优于双流方法。同时提出合成数据增强策略,生成高质量数据,构建出新的基准数据集Syned1,更适用于复杂差异描述任务的训练与评估。

原文摘要 · Abstract (English)

The rise of the generative models quality during the past years enabled the generation of edited variations of images at an important scale. To counter the harmful effects of such technology, the Image Difference Captioning (IDC) task aims to describe the differences between two images. While this task is successfully handled for simple 3D rendered images, it struggles on real-world images. The reason is twofold: the training data-scarcity, and the difficulty to capture fine-grained differences between complex images. To address those issues, we propose in this paper a simple yet effective framework to both adapt existing image captioning models to the IDC task and augment IDC datasets. We introduce BLIP2IDC, an adaptation of BLIP2 to the IDC task at low computational cost, and show it outperforms two-streams approaches by a significant margin on real-world IDC datasets. We also propose to use synthetic augmentation to improve the performance of IDC models in an agnostic fashion. We show that our synthetic augmentation strategy provides high quality data, leading to a challenging new dataset well-suited for IDC named Syned1.

图像差异描述合成数据BLIP2视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。