arXiv:2510.27263cs.LG2025-10ICCV被引 3

构建首个统一的分布外性能预测评估基准,助力模型安全部署。

ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction

  • 设计统一评估框架,整合主流OOD数据集和算法
  • 提供预训练模型,确保实验可复现与公平对比
  • 分析算法能力边界,指导实际应用选择

近年来,分布外(OOD)性能预测受到越来越多关注,其目标是预测已训练模型在未标注的OOD测试数据上的表现,从而更好地在高风险场景中利用现成模型。尽管已有进展,但以往研究的评估协议不一致,且仅覆盖有限的真实世界OOD数据集和分布偏移类型。为此,我们提出分布外性能预测基准(ODP-Bench),涵盖最常见的OOD数据集及现有实用性能预测算法。我们提供预训练模型作为测试平台,确保比较一致性,并避免重复训练负担。此外,还进行深入实验分析,以更清楚理解各算法的能力边界。

原文摘要 · Abstract (English)

Recently, there has been gradually more attention paid to Out-of-Distribution (OOD) performance prediction, whose goal is to predict the performance of trained models on unlabeled OOD test datasets, so that we could better leverage and deploy off-the-shelf trained models in risk-sensitive scenarios. Although progress has been made in this area, evaluation protocols in previous literature are inconsistent, and most works cover only a limited number of real-world OOD datasets and types of distribution shifts. To provide convenient and fair comparisons for various algorithms, we propose Out-of-Distribution Performance Prediction Benchmark (ODP-Bench), a comprehensive benchmark that includes most commonly used OOD datasets and existing practical performance prediction algorithms. We provide our trained models as a testbench for future researchers, thus guaranteeing the consistency of comparison and avoiding the burden of repeating the model training process. Furthermore, we also conduct in-depth experimental analyses to better understand their capability boundary.

OOD检测模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。