大模型时代,预训练知识比无标签数据更有效于图像分类。
Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models
- 对比无标签数据与预训练模型在小标注量下的表现
- 预训练模型在主流基准上显著优于传统自监督学习方法
- 提示需探索两者深度融合的新路径,适合关注模型效率的研究者
自监督学习(SSL)通过利用无标签数据降低标注成本并取得良好效果。随着大型基础模型的发展,利用预训练模型成为缓解下游任务标签稀缺的有力途径,如各类参数高效微调技术。这引发一个关键问题:当标注数据有限时,应依赖无标签数据还是预训练模型?为此,我们在控制标注预算的条件下,对代表性图像分类任务中的SSL方法与预训练模型(如CLIP)进行了公平比较。实验表明,自监督学习在大模型时代已遭遇瓶颈,预训练模型在广泛采用的SSL基准上展现出更高效率与更强性能。这凸显了自监督学习研究者亟需探索新方向,例如将自监督学习与预训练模型进行深度整合。此外,我们还考察了多模态大语言模型(MLLMs)在图像分类任务中的潜力。结果表明,尽管参数规模巨大,MLLMs仍面临显著性能瓶颈,说明即使看似成熟的任务仍极具挑战性。
原文摘要 · Abstract (English)
Semi-supervised learning (SSL) alleviates the cost of data labeling process by exploiting unlabeled data and has achieved promising results. Meanwhile, with the development of large foundation models, exploiting pre-trained models becomes a promising way to address the label scarcity in the downstream tasks, such as various parameter-efficient fine-tuning techniques. This raises a natural yet critical question: When labeled data is limited, should we rely on unlabeled data or pre-trained models? To investigate this issue, we conduct a fair comparison between SSL methods and pre-trained models (e.g., CLIP) on representative image classification tasks under a controlled supervision budget. Experiments reveal that SSL has met its ``Waterloo" in the era of large models, as pre-trained models show both high efficiency and strong performance on widely adopted SSL benchmarks. This underscores the urgent need for SSL researchers to explore new avenues, such as deeper integration between the SSL and pre-trained models. Furthermore, we investigate the potential of Multi-Modal Large Language Models (MLLMs) in image classification tasks. Results show that, despite their massive parameter scales, MLLMs still face significant performance limitations, highlighting that even a seemingly well-studied task remains highly challenging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。