通过对比相似图像差异,提升多模态大模型的精细视觉感知能力。
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- 设计差异定位任务,让模型在无先验条件下识别图像间细微差别。
- 在RefCOCO等基准上显著提升细粒度感知性能,迁移效果良好。
- 适合关注视觉推理与跨模态理解的研究者使用。
多模态大语言模型在多种视觉-语言任务中表现优异,但其细粒度视觉感知和精确空间推理能力仍受限。本文提出DiG(差分定位)新框架,让模型通过识别和定位相似图像对之间的所有差异来学习细粒度感知,无需预先知道差异数量。为支持可扩展训练,我们构建基于3D渲染的自动化数据生成流水线,生成高质量、可控差异的图像对。针对差异信号稀疏问题,引入课程学习策略,从单个差异逐步过渡到多个差异,实现稳定优化。大量实验证明,DiG显著提升模型在多种视觉感知基准上的表现,且习得的细粒度感知能力能有效迁移到下游任务,包括RefCOCO、RefCOCO+、RefCOCOg及通用多模态感知基准。结果表明,差分定位是一种可扩展且鲁棒的提升多模态大模型细粒度视觉推理能力的方法。
原文摘要 · Abstract (English)
Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introduce DiG (Differential Grounding), a novel proxy task framework where MLLMs learn fine-grained perception by identifying and localizing all differences between similar image pairs without prior knowledge of their number. To support scalable training, we develop an automated 3D rendering-based data generation pipeline that produces high-quality paired images with fully controllable discrepancies. To address the sparsity of difference signals, we further employ curriculum learning that progressively increases complexity from single to multiple differences, enabling stable optimization. Extensive experiments demonstrate that DiG significantly improves model performance across a variety of visual perception benchmarks and that the learned fine-grained perception skills transfer effectively to standard downstream tasks, including RefCOCO, RefCOCO+, RefCOCOg, and general multimodal perception benchmarks. Our results highlight differential grounding as a scalable and robust approach for advancing fine-grained visual reasoning in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。