用1000张图就能高效微调大模型,无需标签和复杂调参。
DIET-CP: Lightweight and Data Efficient Self Supervised Continued Pretraining
- 基于简单无监督目标,仅需少量数据即可持续预训练
- 在仅1000张图像下提升DINOv3等顶尖模型性能
- 适合资源有限的垂直领域,兼容多种模型和模态
持续预训练为将基础模型适配新目标领域提供了有前景的解决方案。然而,在专业领域中,可用数据集通常非常小,限制了面向大规模预训练设计的自监督方法的应用,且难以进行超参数搜索。此外,预训练模型通常仅以骨干权重形式发布,缺少继续预训练所需的重要信息。我们提出DIET-CP,一种简单的持续预训练策略,可将任意强基座模型引导至感兴趣的新型数据分布。DIET-CP采用极简目标函数,无需标签,引入的超参数数量与监督微调相当。该方法在不同数据模态和骨干结构间表现稳定,并在仅使用1000张图像的情况下,显著提升DINOv3等先进模型的性能。
原文摘要 · Abstract (English)
Continued pretraining offers a promising solution for adapting foundation models to a new target domain. However, in specialized domains, available datasets are often very small, limiting the applicability of SSL methods developed for large-scale pretraining and making hyperparameter search infeasible. In addition, pretrained models are usually released as backbone-weights only, lacking important information to continue pretraining. We propose to bridge this gap with DIET-CP, a simple continued pretraining strategy, where any strong foundation model can be steered towards the new data distribution of interest. DIET-CP relies on a very simple objective, requires no labels, and introduces no more hyperparameters than supervised finetuning. It is stable across data modalities and backbone choices, while providing a significant performance boost for state-of-the-art models such as DINOv3 using only 1000 images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。