arXiv:2506.04209cs.CV2025-06被引 5

用固定大语言模型做文本编码,仅训练图像编码器就能实现高效图文对齐

Language-Image Alignment with Fixed Text Encoders

  • 仅训练图像编码器,使用预训练固定语言模型作为文本编码器
  • 在组合理解与长描述任务中超越CLIP,计算开销显著降低
  • 为大模型驱动视觉学习提供新范式,适合追求效率的视觉-语言研究者

当前主流的图文对齐方法通过对比学习联合预训练文本和图像编码器,如CLIP及其变体。本文质疑这种昂贵的联合训练是否必要。我们探究了预训练的大语言模型(LLM)能否作为足够好的文本编码器来引导视觉表征学习。为此,提出仅训练图像编码器的固定文本编码器图文对齐框架LIFT。通过全面的基准测试与消融实验发现,该简化框架LIFT在多数涉及组合理解与长描述的任务中表现优于CLIP,同时在计算效率上取得显著提升。本工作首次系统探索了来自大语言模型的文本嵌入如何引导视觉学习,为构建语言对齐的视觉表征提供了替代设计选择。

原文摘要 · Abstract (English)

Currently, the most dominant approach to establishing language-image alignment is to pre-train text and image encoders jointly through contrastive learning, such as CLIP and its variants. In this work, we question whether such a costly joint training is necessary. In particular, we investigate if a pre-trained fixed large language model (LLM) offers a good enough text encoder to guide visual representation learning. That is, we propose to learn Language-Image alignment with a Fixed Text encoder (LIFT) from an LLM by training only the image encoder. Somewhat surprisingly, through comprehensive benchmarking and ablation studies, we find that this much simplified framework LIFT is highly effective and it outperforms CLIP in most scenarios that involve compositional understanding and long captions, while achieving considerable gains in computational efficiency. Our work takes a first step towards systematically exploring how text embeddings from LLMs can guide visual learning and suggests an alternative design choice for learning language-aligned visual representations.

图文对齐大语言模型效率优化视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。