arXiv:2507.04699cs.CV2025-07ICCV被引 5

用生成反事实图像对提升视觉语言模型的组合推理能力

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

  • 通过扩散模型生成带空间关系的图像块,自动构造反事实数据集
  • 在多个基准上显著提升视觉推理性能,训练数据量减少超50%
  • 适合研究视觉语言模型、数据增强与生成式训练的学者

视觉语言模型常因高质量图文数据不足而难以进行组合推理。为此,我们提出一种基于块的扩散方法,无需人工标注即可自动生成反事实数据集。该方法利用大语言模型识别实体及其空间关系,独立生成符合组合规则的图像块,实现高保真、多样化的反事实图文对。同时引入专用损失函数,区分集间与集内样本,提升训练效率并减少负样本需求。实验表明,用该反事实数据微调视觉语言模型,在多个基准上取得领先性能,且所需训练数据远少于现有方法。

原文摘要 · Abstract (English)

Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes large language models to identify entities and their spatial relationships. It then independently generates image blocks as "puzzle pieces" coherently arranged according to specified compositional rules. This process creates diverse, high-fidelity counterfactual image-text pairs with precisely controlled variations. In addition, we introduce a specialized loss function that differentiates inter-set from intra-set samples, enhancing training efficiency and reducing the need for negative samples. Experiments demonstrate that fine-tuning VLMs with our counterfactual datasets significantly improves visual reasoning performance. Our approach achieves state-of-the-art results across multiple benchmarks while using substantially less training data than existing methods.

视觉语言模型生成对抗组合推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。