arXiv:2505.18983cs.LGcs.CV2025-05NeurIPS被引 2

用轻量网络摊销对比学习计算,显著提升CLIP训练效率和性能。

AmorLIP: Efficient Language-Image Pretraining via Amortization

  • 用轻量神经网络摊销对比学习中的昂贵计算
  • 在38个下游任务上零样本性能提升最高达12.24%
  • 适合追求高效预训练的视觉语言研究者

对比语言-图像预训练(CLIP)在多种下游文本-图像任务中展现出强大的零样本性能。现有方法通常使用每个小批量中的负样本优化对比目标,为实现稳健的表示学习,这些方法需要极大的批量大小,导致计算需求上升至数百甚至上千块GPU。此前的缓解方案往往牺牲下游性能、延长训练时间,或在超大数据集下面临可扩展性挑战。为此,我们提出AmorLIP,一种高效的CLIP预训练框架,通过轻量神经网络摊销对比学习中的高成本计算,显著提升训练效率与性能。基于能量模型的谱分解洞察,我们引入新的摊销目标及实用技术以增强训练稳定性。在38个下游任务上的广泛实验表明,AmorLIP具备更优的零样本分类与检索能力,相较标准CLIP基线有高达12.24%的相对提升。

原文摘要 · Abstract (English)

Contrastive Language-Image Pretraining (CLIP) has demonstrated strong zero-shot performance across diverse downstream text-image tasks. Existing CLIP methods typically optimize a contrastive objective using negative samples drawn from each minibatch. To achieve robust representation learning, these methods require extremely large batch sizes and escalate computational demands to hundreds or even thousands of GPUs. Prior approaches to mitigate this issue often compromise downstream performance, prolong training duration, or face scalability challenges with very large datasets. To overcome these limitations, we propose AmorLIP, an efficient CLIP pretraining framework that amortizes expensive computations involved in contrastive learning through lightweight neural networks, which substantially improves training efficiency and performance. Leveraging insights from a spectral factorization of energy-based models, we introduce novel amortization objectives along with practical techniques to improve training stability. Extensive experiments across 38 downstream tasks demonstrate the superior zero-shot classification and retrieval capabilities of AmorLIP, consistently outperforming standard CLIP baselines with substantial relative improvements of up to 12.24%.

CLIP对比学习高效训练语言图像预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。