arXiv:2510.06448cs.LGcs.AI2025-10被引 2

现有模型迁移性能评估方法存在严重缺陷,真实场景下效果大打折扣。

How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation

  • 指出当前评估体系的模型空间和性能排名过于静态,脱离实际。
  • 发现简单启发式方法在虚假评测中反而超过复杂算法。
  • 呼吁构建更贴近真实场景的动态评估基准,指导未来研究。

迁移能力评估指标用于在不微调模型且无源数据访问的情况下,为特定目标任务选择表现优异的预训练模型。尽管此类指标备受关注,其评估基准却长期未受审视。本文通过实证表明,广泛使用的基准设置存在根本性缺陷。其不合理的模型空间与静态性能排序人为夸大了现有指标的表现,导致简单的、与数据集无关的启发式方法反而优于复杂方法。分析揭示了当前评估协议与真实模型选择复杂性的严重脱节。为此,本文提出具体建议,以构建更稳健、更真实的评估基准,引导未来研究走向更具意义的方向。

原文摘要 · Abstract (English)

Transferability estimation metrics are used to find a high-performing pre-trained model for a given target task without fine-tuning models and without access to the source dataset. Despite the growing interest in developing such metrics, the benchmarks used to measure their progress have gone largely unexamined. In this work, we empirically show the shortcomings of widely used benchmark setups to evaluate transferability estimation metrics. We argue that the benchmarks on which these metrics are evaluated are fundamentally flawed. We empirically demonstrate that their unrealistic model spaces and static performance hierarchies artificially inflate the perceived performance of existing metrics, to the point where simple, dataset-agnostic heuristics can outperform sophisticated methods. Our analysis reveals a critical disconnect between current evaluation protocols and the complexities of real-world model selection. To address this, we provide concrete recommendations for constructing more robust and realistic benchmarks to guide future research in a more meaningful direction.

模型评估迁移学习基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。