提出标准化框架,公平评估模型跨域迁移能力。
Benchmarking Transferability: A Framework for Fair and Robust Evaluation
- 构建统一评估框架,消除不同实验设置干扰。
- 在头训练微调中,新指标提升3.5%性能。
- 适合需要可靠迁移评估的研究者使用。
迁移性评分旨在量化模型在源域训练后对目标域的泛化能力。尽管已有多种方法用于衡量迁移性,但其可靠性与实际效用仍不明确,常因实验设置、数据集和假设差异而产生偏差。本文提出一个全面的基准测试框架,系统评估不同场景下的迁移性评分。通过大量实验发现,各类指标在不同情况下表现各异,表明现有评估方式可能未能充分揭示方法的优势与局限。研究强调标准化评估协议的重要性,为更可靠的迁移性度量及跨域应用中的模型选择提供支持。此外,在头训练微调实验设置中,新提出的指标实现3.5%的性能提升。代码已开源:https://github.com/alizkzm/pert_robust_platform。
原文摘要 · Abstract (English)
Transferability scores aim to quantify how well a model trained on one domain generalizes to a target domain. Despite numerous methods proposed for measuring transferability, their reliability and practical usefulness remain inconclusive, often due to differing experimental setups, datasets, and assumptions. In this paper, we introduce a comprehensive benchmarking framework designed to systematically evaluate transferability scores across diverse settings. Through extensive experiments, we observe variations in how different metrics perform under various scenarios, suggesting that current evaluation practices may not fully capture each method's strengths and limitations. Our findings underscore the value of standardized assessment protocols, paving the way for more reliable transferability measures and better-informed model selection in cross-domain applications. Additionally, we achieved a 3.5\% improvement using our proposed metric for the head-training fine-tuning experimental setup. Our code is available in this repository: https://github.com/alizkzm/pert_robust_platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。