大模型语言能力可弥补视觉特征不足,提升多模态表现。
LLMs Can Compensate for Deficiencies in Visual Representations
- 用消融实验验证语言模型能补足弱视觉特征
- 视觉信息缺失时,语言解码器仍可恢复性能
- 适合研究多模态架构设计与语言补偿机制的学者
许多在多模态任务中表现优异的视觉-语言模型(VLMs)基于CLIP视觉编码器,而该编码器存在已知局限。本文提出假设:强大的语言主干可通过上下文化或丰富视觉特征来弥补其不足。通过在三个基于CLIP的VLM上进行受控自注意力消融实验,结果表明尽管存在缺陷,CLIP的视觉表示仍包含可直接读取的语义信息。当视觉表示上下文信息减少时,语言解码器能显著补偿视觉缺陷并恢复性能。这揭示了多模态模型中的动态分工机制,为未来将更多视觉处理任务交由语言解码器的设计提供了依据。
原文摘要 · Abstract (English)
Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We investigate the hypothesis that the strong language backbone in VLMs compensates for possibly weak visual features by contextualizing or enriching them. Using three CLIP-based VLMs, we perform controlled self-attention ablations on a carefully designed probing task. Our findings show that despite known limitations, CLIP visual representations offer ready-to-read semantic information to the language decoder. However, in scenarios of reduced contextualization in the visual representations, the language decoder can largely compensate for the deficiency and recover performance. This suggests a dynamic division of labor in VLMs and motivates future architectures that offload more visual processing to the language decoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。