用低成本方法提升大模型对多图细节的理解能力
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding

- 通过图像间与图像内对比构建多图训练数据
- 在多个基准上达到当前最佳效果,最高提升2.90分
- 适合需要精准多图理解的视觉推理任务
尽管多模态大语言模型发展迅速,但在细粒度多图理解方面仍存在空间幻觉、注意力泄露和物体一致性失败等问题。现有方法通常依赖昂贵的人工标注或大规模思维链数据生成。本文提出组合式基底对比(CGC),一种低成本全框架,用于增强多模态大模型的细粒度多图理解能力。基于已有单图定位标注,CGC通过图像间对比引入语义解耦的干扰上下文以实现跨图区分,通过图像内对比构建相关跨视角样本以维持物体一致性。同时,在GRPO框架中引入基于规则的空间奖励机制,提升源图归属、空间对齐和结构化输出有效性,遵循‘先思考后定位’范式。实验表明,CGC在细粒度多图基准MIG-Bench和VLM2-Bench上达到最优表现;所学能力还迁移至更广泛的多模态理解与推理任务,在MathVista(+2.90)、MuirBench(+2.88)、MMStar(+1.93)、MMMU(+1.77)和BLINK(+1.69)上均优于Qwen3-VL-8B基础模型。
原文摘要 · Abstract (English)
Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding, often exhibiting spatial hallucination, attention leakage, and failures in object constancy. In addition, existing approaches typically rely on expensive human annotations or large-scale chain-of-thought (CoT) data generation. We propose Compositional Grounded Contrast (abbr. CGC), a low-cost full framework for boosting fine-grained multi-image understanding of MLLMs. Built on existing single-image grounding annotations, CGC constructs compositional multi-image training instances through Inter-Image Contrast and Intra-Image Contrast, which introduce semantically decoupled distractor contexts for cross-image discrimination and correlated cross-view samples for object constancy, respectively. CGC further introduces a Rule-Based Spatial Reward within the GRPO framework to improve source-image attribution, spatial alignment, and structured output validity under a Think-before-Grounding paradigm. Experiments show that CGC achieves state-of-the-art results on fine-grained multi-image benchmarks, including MIG-Bench and VLM2-Bench. The learned multi-image understanding capability also transfers to broader multimodal understanding and reasoning tasks, yielding consistent gains over the Qwen3-VL-8B base model on MathVista (+2.90), MuirBench (+2.88), MMStar (+1.93), MMMU (+1.77), and BLINK (+1.69).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。