arXiv:2502.06925cs.LGcs.AI2025-02

提出更简单模型评估迁移能力,提升选型准确率

Occam's model: Selecting simpler representations for better transferability estimation

  • 用模型表示的简洁性衡量迁移潜力
  • 相比顶尖方法,相关性提升最高32%
  • 适合模型选型与跨任务迁移研究者

在大规模数据上预训练的模型微调已成为现代机器学习的核心流程。随着Hugging Face等在线模型库的普及,为特定任务微调预训练模型变得前所未有的便捷。这引出一个关键问题:哪个预训练模型最适合当前任务?该问题称为迁移能力估计。本文提出两种新颖且有效的迁移能力评估指标。我们的方法将迁移能力视为预训练模型表示在多分类任务中被快速区分的难易程度,提供了全新的视角。我们在多种任务设置下严格评估所提指标,相较于现有最优方法展现出更强的鲁棒性和实用性。此外,我们提供理论分析,解释指标的有效性与泛化能力。实验表明,所提方法在Kendall's Tau相关性上最高提升32%。

原文摘要 · Abstract (English)

Fine-tuning models that have been pre-trained on large datasets has become a cornerstone of modern machine learning workflows. With the widespread availability of online model repositories, such as Hugging Face, it is now easier than ever to fine-tune pre-trained models for specific tasks. This raises a critical question: which pre-trained model is most suitable for a given task? This problem is called transferability estimation. In this work, we introduce two novel and effective metrics for estimating the transferability of pre-trained models. Our approach is grounded in viewing transferability as a measure of how easily a pre-trained model's representations can be trained to separate target classes, providing a unique perspective on transferability estimation. We rigorously evaluate the proposed metrics against state-of-the-art alternatives across diverse problem settings, demonstrating their robustness and practical utility. Additionally, we present theoretical insights that explain our metrics' efficacy and adaptability to various scenarios. We experimentally show that our metrics increase Kendall's Tau by up to 32% compared to the state-of-the-art baselines.

模型选型迁移学习评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。