通过内容风格分解提升视觉模型在少量标注数据下的表现
Semi-Supervised Fine-Tuning of Vision Foundation Models with Content-Style Decomposition
- 基于信息论框架实现内容与风格的解耦表征
- 在低标注数据下显著优于监督微调基线,尤其在冻结和可训练主干上
- 适用于数据稀缺场景,对迁移学习研究者有参考价值
本文提出一种半监督微调方法,旨在提升预训练基础模型在下游任务中有限标注数据下的性能。通过在信息论框架内引入内容-风格分解,该方法增强预训练视觉基础模型的隐含表示,使其更契合特定任务目标,并缓解分布偏移问题。我们在多个数据集上进行了评估,包括MNIST及其带黄色和白色条纹的增强版本、CIFAR-10、SVHN和GalaxyMNIST。实验结果表明,在大多数测试数据集上,该方法在低标注数据环境下均优于监督微调基线,无论采用冻结或可训练主干结构均有效。
原文摘要 · Abstract (English)
In this paper, we present a semi-supervised fine-tuning approach designed to improve the performance of pre-trained foundation models on downstream tasks with limited labeled data. By leveraging content-style decomposition within an information-theoretic framework, our method enhances the latent representations of pre-trained vision foundation models, aligning them more effectively with specific task objectives and addressing the problem of distribution shift. We evaluate our approach on multiple datasets, including MNIST, its augmented variations (with yellow and white stripes), CIFAR-10, SVHN, and GalaxyMNIST. The experiments show improvements over supervised finetuning baseline of pre-trained models, particularly in low-labeled data regimes, across both frozen and trainable backbones for the majority of the tested datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。