让大模型更懂图像,通过智能融合视觉与语言特征提升识别能力
CARPE: Context-Aware Image Representation Prioritization via Ensemble for Large Vision-Language Models

- 用上下文感知的集成机制融合原始视觉特征和语言对齐特征
- 在图像分类和多任务基准上显著优于原模型
- 适合需要强化视觉理解的多模态应用开发者
大型视觉语言模型(LVLM)通常采用自回归语言建模目标进行训练,将视觉表征与语言空间对齐。尽管该方法在多模态推理中表现良好,但会削弱以视觉为中心的能力,导致其在图像分类等任务上表现不如基础视觉编码器。为此,我们提出一种轻量级框架CARPE(Context-Aware Image Representation Prioritization via Ensemble),通过视觉融合层与上下文感知集成机制,将原始视觉特征与对齐的LLM表征相结合。该设计使模型能自适应地权衡视觉与文本模态,捕捉图像表征的多个层面。大量实验表明,CARPE在图像分类及多种视觉语言基准上均实现性能提升。结果表明,在自回归型LVLM中,模态平衡对多模态泛化至关重要,可有效提升表征利用效率。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) are typically trained using autoregressive language modeling objectives, which align visual representations with linguistic space. While effective for multimodal reasoning, this alignment can weaken vision-centric capabilities, causing LVLMs to underperform their base vision encoders on tasks such as image classification. To address this limitation, we propose Context-Aware Image Representation Prioritization via Ensemble (CARPE), a lightweight framework that integrates raw vision features with aligned LLM representations through vision-integration layers and a context-aware ensemble mechanism. This design enhances the model's ability to adaptively weight visual and textual modalities and enables the model to capture various aspects of image representations. Extensive experiments demonstrate that CARPE improves performance on both image classification and diverse vision-language benchmarks. Our results suggest that modality balancing plays a critical role in multimodal generalization by improving representation utilization within autoregressive LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。