arXiv:2503.18034cs.CVcs.CL2025-03Conference of the …被引 1

提升视觉模型先验知识,可显著增强多模态大模型的视觉理解能力。

Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models

  • 提出新指标Rank_e量化视觉编码器先验知识对性能的影响。
  • 发现先验知识越强,多模态模型表现越好,尤其在罕见视觉实体上。
  • 设计两阶段框架VisPRE,主动增强视觉编码器的先验知识,适合低频视觉任务。

现有研究大多将多模态大语言模型(MLLM)视为统一系统进行端到端训练,却很少关注视觉编码器先验知识的影响。本文提出新指标Rank_e,量化视觉编码器先验知识对MLLM性能的影响。分析表明,先验知识与模型表现呈正相关。此外,仅用端到端视觉问答(VQA)数据进行领域特定微调,对缺乏固有视觉先验的知识实体效果有限。为此,我们提出VisPRE(Vision Prior Remediation)——一种两阶段训练框架,通过在视觉编码器层面显式引入先验知识。实验结果表明,增强视觉编码器的先验知识能显著提升MLLM的视觉理解能力,尤其在涉及不常见视觉实体的场景中,为性能优化提供了新颖有效策略。

原文摘要 · Abstract (English)

Does the prior knowledge of the vision encoder constrain the capability boundary of Multi-modal Large Language Models (MLLMs)? While most existing research treats MLLMs as unified systems optimized through end-to-end training, the impact of vision encoder's prior knowledge is seldom investigated. In this work, we introduce a novel metric, $Rank_e$, to quantify the effect of prior knowledge of the vision encoder on MLLM performance. Our analysis reveals a positive correlation between prior knowledge and MLLM performance. Moreover, we find that domain-specific fine-tuning using solely end-to-end visual question answering (VQA) data is insufficient, particularly for entities with low inherent visual prior knowledge. To address this issue, we propose VisPRE (Vision Prior Remediation), a two-stage training framework that explicitly incorporates prior knowledge at the vision encoder level. Experimental results demonstrate that augmenting vision encoder's prior knowledge substantially boosts the visual understanding capabilities of MLLMs, offering a novel and effective strategy for improving performance, especially in scenarios involving uncommon visual entities.

多模态视觉先验模型优化知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。