通过视觉认知对齐提升图文模型理解能力
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
- 构建多粒度地标数据集,分析视觉编码器未知数据的影响
- 提出EECA方法,用多粒度监督生成对齐语言框架的视觉令牌
- 在地标识别任务中显著提升模型性能,适合多模态系统研究者
大型视觉语言模型(LVLM)将独立预训练的视觉与语言组件结合,常以CLIP-ViT作为视觉主干。然而,这类模型普遍存在视觉编码器(VE)与大语言模型(LLM)之间的“认知错位”问题:视觉表示可能超出语言模型的认知范围,导致理解偏差。本文研究不同视觉表示对模型理解力的影响,尤其关注当语言模型面对视觉编码器未知(VE-Unknown)图像时,其模糊表示如何挑战解释精度。为此,我们构建了多粒度地标数据集,系统评估了VE-Known与VE-Unknown数据对解释能力的影响。结果表明,VE-Unknown数据会限制模型准确理解的能力,而富含显著特征的VE-Known数据有助于缓解认知错位。基于此,我们提出实体增强认知对齐(EECA)方法,通过多粒度监督生成视觉丰富且与语言模型认知框架对齐的令牌,不仅嵌入于语言模型的嵌入空间,更提升整体一致性。该方法显著增强了在地标识别任务中的表现,凸显了认知对齐在多模态系统中的关键作用。
原文摘要 · Abstract (English)
Does seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. However, these models frequently encounter a core issue of "cognitive misalignment" between the vision encoder (VE) and the large language model (LLM). Specifically, the VE's representation of visual information may not fully align with LLM's cognitive framework, leading to a mismatch where visual features exceed the language model's interpretive range. To address this, we investigate how variations in VE representations influence LVLM comprehension, especially when the LLM faces VE-Unknown data-images whose ambiguous visual representations challenge the VE's interpretive precision. Accordingly, we construct a multi-granularity landmark dataset and systematically examine the impact of VE-Known and VE-Unknown data on interpretive abilities. Our results show that VE-Unknown data limits LVLM's capacity for accurate understanding, while VE-Known data, rich in distinctive features, helps reduce cognitive misalignment. Building on these insights, we propose Entity-Enhanced Cognitive Alignment (EECA), a method that employs multi-granularity supervision to generate visually enriched, well-aligned tokens that not only integrate within the LLM's embedding space but also align with the LLM's cognitive framework. This alignment markedly enhances LVLM performance in landmark recognition. Our findings underscore the challenges posed by VE-Unknown data and highlight the essential role of cognitive alignment in advancing multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。