arXiv:2503.07851cs.LGcs.CV2025-03

用互信息分解提升小样本下的模型微调效果

TwinTURBO: Semi-Supervised Fine-Tuning of Foundation Models via Mutual Information Decompositions for Downstream Task and Latent Spaces

  • 通过分解互信息,分别优化任务空间和隐空间表示
  • 在极低标注数据下,分类性能显著提升
  • 适合资源受限场景下微调大模型的科研与工程人员

我们提出一种针对基础模型的半监督微调框架,利用互信息分解应对标注数据有限的挑战。该方法推导出两个独立的下界:其一用于下游任务空间(如分类),通过条件与边缘交叉熵及KL散度优化;其二用于隐空间表示,采用类对比分解进行正则化与对齐。该微调策略保留预训练模型结构,仅修改一个包含小型Transformer和标记聚合技术的专用投影模块。在多个数据集上的实验表明,在极端低标注条件下,该方法能有效利用未标注数据,显著提升分类性能。

原文摘要 · Abstract (English)

We present a semi-supervised fine-tuning framework for foundation models that utilises mutual information decomposition to address the challenges of training for a limited amount of labelled data. Our approach derives two distinct lower bounds: i) for the downstream task space, such as classification, optimised using conditional and marginal cross-entropy alongside Kullback-Leibler divergence, and ii) for the latent space representation, regularised and aligned using a contrastive-like decomposition. This fine-tuning strategy retains the pre-trained structure of the foundation model, modifying only a specialised projector module comprising a small transformer and a token aggregation technique. Experiments on several datasets demonstrate significant improvements in classification tasks under extremely low-labelled conditions by effectively leveraging unlabelled data.

半监督学习模型微调互信息小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。