arXiv:2508.06525cs.CV2025-08

大模型能通过语言推理提升图像识别准确率,且只需少量文本信息即可。

Large Language Models Facilitate Vision Reflection in Image Classification

  • 用大模型验证专用视觉模型预测,提升图像分类准确率。
  • 仅用少量文本令牌就能生成相似回答,说明依赖精炼语义表示。
  • 无需训练即可增强细粒度识别能力,适合追求可解释性的场景。

本文揭示了大型多模态模型(LMMs)在视觉反思中的若干新发现。首先,我们证明,通过提示LMM验证专用视觉模型的预测,即使在ImageNet等基准上也能提升识别准确率,尽管此前研究表明LMM通常不如专用视觉编码器。其次,我们分析了视觉反思的内部行为,发现视觉-语言连接器将视觉特征映射为明确的文本概念,使语言模型能够利用常识知识判断预测合理性。我们进一步观察到,用少量文本令牌替换大量视觉令牌后,LLaVA仍能生成相似答案,表明LMM可能主要依赖一组精炼的文本表示而非原始视觉特征。第三,我们展示了一种无需训练的连接器可在细粒度识别任务中提升性能,无需复杂的特征对齐训练。这些发现为视觉-语言模型的可解释性提供了新视角,并表明视觉反思是实现鲁棒且可解释视觉识别的可行策略。

原文摘要 · Abstract (English)

This paper presents several novel findings on the explainability of vision reflection in large multimodal models (LMMs). First, we show that prompting an LMM to verify the prediction of a specialized vision model can improve recognition accuracy, even on benchmarks like ImageNet, despite prior evidence that LMMs typically underperform dedicated vision encoders. Second, we analyze the internal behavior of vision reflection and find that the vision-language connector maps visual features into explicit textual concepts, allowing the language model to reason about prediction plausibility using commonsense knowledge. We further observe that replacing a large number of vision tokens with only a few text tokens still enables LLaVA to generate similar answers, suggesting that LMMs may rely primarily on a compact set of distilled textual representations rather than raw vision features. Third, we show that a training-free connector can enhance LMM performance in fine-grained recognition tasks, without extensive feature-alignment training. Together, these findings offer new insights into the explainability of vision-language models and suggest that vision reflection is a promising strategy for achieving robust and interpretable visual recognition.

多模态模型视觉反思可解释性图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。