arXiv:2508.08066cs.CVcs.AI2025-08被引 2

系统研究多模态大模型视觉定位设计,提升精准度。

ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model

  • 对比多种视觉定位方法,找出最优设计。
  • 通过数据优化使模型在三个数据集上提升5.6%~7.0%。
  • 适合想改进视觉定位性能的研究者参考。

多模态大模型(MLLM)的细粒度多模态能力已成为关键研究方向,尤其在解决视觉定位(VG)问题方面。尽管现有方法表现良好,但其在微调MLLM进行视觉定位时采用的设计选择各异,缺乏系统性验证。为此,本文对影响MLLM视觉定位性能的各种设计选择进行了全面研究。分析基于广泛采用的LLaVA-1.5模型展开,虽存在更新模型,但为确保结论普适性与可扩展性,仍沿用此标准。研究涵盖两个核心方面:(1) 探索不同视觉定位范式,识别最有效设计并提供见解;(2) 对定位数据的设计进行消融实验,优化模型微调策略。最终成果显著提升了模型在视觉定位任务上的表现,在RefCOCO+/+g三个数据集上分别取得+5.6%、+6.9%、+7.0%的性能提升。

原文摘要 · Abstract (English)

Fine-grained multimodal capability in Multimodal Large Language Models (MLLMs) has emerged as a critical research direction, particularly for tackling the visual grounding (VG) problem. Despite the strong performance achieved by existing approaches, they often employ disparate design choices when fine-tuning MLLMs for VG, lacking systematic verification to support these designs. To bridge this gap, this paper presents a comprehensive study of various design choices that impact the VG performance of MLLMs. We conduct our analysis using LLaVA-1.5, which has been widely adopted in prior empirical studies of MLLMs. While more recent models exist, we follow this convention to ensure our findings remain broadly applicable and extendable to other architectures. We cover two key aspects: (1) exploring different visual grounding paradigms in MLLMs, identifying the most effective design, and providing our insights; and (2) conducting ablation studies on the design of grounding data to optimize MLLMs' fine-tuning for the VG task. Finally, our findings contribute to a stronger MLLM for VG, achieving improvements of +5.6% / +6.9% / +7.0% on RefCOCO/+/g over the LLaVA-1.5.

视觉定位多模态模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。