arXiv:2505.05626cs.CVcs.AI2025-05被引 5

提升视觉理解,让模型不再依赖语言先验。

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models

  • 分析模型内部视觉认知机制,优化视觉信息提取
  • 在视觉挑战任务上性能提升10个百分点
  • 适合需要精准视觉推理的多模态应用

实现视觉与语言的深度对齐仍是多模态大语言模型(MLLMs)的核心挑战。这些模型常无法充分利用视觉输入,倾向于依赖强语言先验。本文首先揭示了MLLMs如何内部构建图像区域的视觉理解,随后提出增强该能力的技术。具体而言,我们设计方法以深化模型对视觉内容的理解,并确保这些视觉洞察能主动引导语言生成。通过详尽的上游分析,验证了所提模型在预测依赖视觉的词元上的优势,以及在视觉挑战任务上取得10个百分点的性能提升。

原文摘要 · Abstract (English)

Achieving deep alignment between vision and language remains a central challenge for Multimodal Large Language Models (MLLMs). These models often fail to fully leverage visual input, defaulting to strong language priors. Our approach first provides insights into how MLLMs internally build visual understanding of image regions and then introduces techniques to amplify this capability. Specifically, we explore techniques designed both to deepen the model's understanding of visual content and to ensure that these visual insights actively guide language generation. We demonstrate the superior multimodal understanding of our resultant model through a detailed upstream analysis quantifying its ability to predict visually-dependent tokens as well as 10 pt boost on visually challenging tasks.

多模态视觉理解语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。