用3000万数据实现顶尖性能,平衡监督信号提升图文模型效率。
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
- 通过自蒸馏与不确定性加权,智能融合多类监督信号。
- 仅用3000万图像,多项任务超越同类高效模型。
- 适合追求低资源高性能的视觉语言研究者使用。
以往大规模视觉语言模型的成功主要依赖于数十亿样本数据集,成为发展瓶颈。近期工作尝试提升监督质量以缩小差距,但仅解决对比学习中的部分问题。本文提出GoldiCLIP框架,基于‘恰到好处’原则,平衡多种监督信号。其多方面训练机制协同包含:(1) 文本条件自蒸馏,对齐无文本与有文本特征;(2) 集成解码器的视觉问答目标,使编码器泛化至非描述性查询;(3) 不确定性加权机制,自动调节异构损失。仅在3000万图像上训练,性能优于现有高效方法,在MSCOCO检索任务上提升2.2点,细粒度检索提升2.0点,基于问题的检索提升5.9点,同时媲美百亿级数据模型。项目主页:https://petsi.uk/goldiclip。
原文摘要 · Abstract (English)
Until recently, the success of large-scale vision-language models (VLMs) has primarily relied on billion-sample datasets, posing a significant barrier to progress. Latest works have begun to close this gap by improving supervision quality, but each addresses only a subset of the weaknesses in contrastive pretraining. We present GoldiCLIP, a framework built on a Goldilocks principle of finding the right balance of supervision signals. Our multifaceted training framework synergistically combines three key innovations: (1) a text-conditioned self-distillation method to align both text-agnostic and text-conditioned features; (2) an encoder integrated decoder with Visual Question Answering (VQA) objective that enables the encoder to generalize beyond the caption-like queries; and (3) an uncertainty-based weighting mechanism that automatically balances all heterogeneous losses. Trained on just 30 million images, 300x less data than leading methods, GoldiCLIP achieves state-of-the-art among data-efficient approaches, improving over the best comparable baseline by 2.2 points on MSCOCO retrieval, 2.0 on fine-grained retrieval, and 5.9 on question-based retrieval, while remaining competitive with billion-scale models. Project page: https://petsi.uk/goldiclip.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。