arXiv:2509.07979cs.CV2025-09被引 37

让多模态大模型更懂视觉细节,提升物体计数与空间推理能力

Visual Representation Alignment for Multimodal Large Language Models

  • 通过对齐预训练视觉模型的内部表征,强化视觉路径学习
  • 在多个基准上实现一致性能提升,尤其改善视觉密集任务表现
  • 适合关注视觉理解、模型训练优化的研究者与开发者

多模态大语言模型(MLLM)在视觉指令微调后已在多种任务中表现强劲,但在物体计数、空间推理等视觉主导任务上仍有局限。我们归因于当前以文本为主的监督范式,仅提供间接视觉路径指导,导致模型在训练中丢失细粒度视觉信息。本文提出视觉表征对齐(VIRAL),一种简单而有效的正则化策略,将MLLM的内部视觉表征与预训练视觉基础模型(VFMs)对齐。通过显式强制对齐,VIRAL使模型既能保留输入视觉编码器中的关键视觉细节,又能从VFMs中补充额外视觉知识,从而增强对复杂视觉输入的推理能力。实验表明,在广泛使用的多模态基准上,所有任务均取得一致改进。我们还进行了全面消融研究,验证了框架核心设计的有效性。这一简单发现为有效整合视觉信息提供了重要方向。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tasks such as object counting or spatial reasoning. We attribute this gap to the prevailing text-only supervision paradigm, which provides only indirect guidance for the visual pathway and often leads MLLMs to discard fine-grained visual details during training. In this paper, we present VIsual Representation ALignment (VIRAL), a simple yet effective regularization strategy that aligns the internal visual representations of MLLMs with those of pre-trained vision foundation models (VFMs). By explicitly enforcing this alignment, VIRAL enables the model not only to retain critical visual details from the input vision encoder but also to complement additional visual knowledge from VFMs, thereby enhancing its ability to reason over complex visual inputs. Our experiments demonstrate consistent improvements across all tasks on widely adopted multimodal benchmarks. Furthermore, we conduct comprehensive ablation studies to validate the key design choices underlying our framework. We believe this simple finding opens up an important direction for the effective integration of visual information in training MLLMs.

多模态视觉对齐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。