arXiv:2504.01017cs.CV2025-04ICCV被引 67

纯视觉自监督学习在大规模下可达到CLIP水平,无需语言监督

Scaling Language-Free Visual Representation Learning

  • 在相同数据上训练视觉自监督与CLIP模型,对比性能
  • 视觉自监督模型在70亿参数下仍持续提升,未达饱和
  • 适合关注视觉主导表征学习的研究者和工程师

当前视觉自监督学习(Visual SSL)在多模态任务如视觉问答(VQA)中表现逊于对比语言图像预训练(CLIP)。这一差距常归因于语言监督带来的语义信息,尽管两者通常使用不同数据训练。本文通过在相同MetaCLIP数据上训练视觉SSL与CLIP模型,并以VQA为多样化评估基准,检验该问题。结果表明,在控制条件下,视觉SSL模型在数据量和模型容量扩大时表现优于CLIP,且在70亿参数规模下仍无性能饱和。因此,视觉自监督方法在多种VQA及经典视觉基准上达到了与CLIP相当的水平。这说明纯视觉自监督学习在规模化后可媲美语言监督的视觉预训练,为以视觉为中心的表示学习开辟新路径。

原文摘要 · Abstract (English)

Visual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This multimodal gap is often attributed to the semantics introduced by language supervision, even though visual SSL and CLIP models are often trained on different data. In this work, we ask the question: "Do visual self-supervised approaches lag behind CLIP due to the lack of language supervision, or differences in the training data?" We study this question by training both visual SSL and CLIP models on the same MetaCLIP data, and leveraging VQA as a diverse testbed for vision encoders. In this controlled setup, visual SSL models scale better than CLIP models in terms of data and model capacity, and visual SSL performance does not saturate even after scaling up to 7B parameters. Consequently, we observe visual SSL methods achieve CLIP-level performance on a wide range of VQA and classic vision benchmarks. These findings demonstrate that pure visual SSL can match language-supervised visual pretraining at scale, opening new opportunities for vision-centric representation learning.

视觉表征自监督学习多模态模型规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。