arXiv:2501.04234stat.MLcs.LG2025-01被引 12

为机器学习基准的综合性能指标提供不确定性量化方法

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

  • 采用自助法与分层贝叶斯建模,量化多任务平均性能的统计不确定性
  • 在VTAB基准上发现:某模型虽整体表现一般,但在特定任务类型中占主导
  • 适合关注模型真实性能差异的研究者和评测工程师

现代人工智能依赖于在海量数据上预训练并适应多种下游任务的机器学习模型(如基础模型)。为总结多任务上的性能表现,常将评估指标聚合为综合指标,例如10个问答任务的平均准确率。在聚合评估指标时,引入综合指标的不确定性有助于更真实地理解模型性能。本文旨在展示如何运用统计方法对跨多任务聚合的性能指标进行不确定性量化。重点方法包括自助法(bootstrapping)、贝叶斯分层(即多层次)建模,以及考虑标准误的任务权重可视化。这些技术揭示了某些模型尽管整体表现不佳,却在特定任务类型中占据主导地位等深层洞察。我们以流行的机器学习基准视觉任务适应基准(Visual Task Adaptation Benchmark, VTAB)为例,验证了所提方法的有效性。

原文摘要 · Abstract (English)

Modern artificial intelligence is supported by machine learning models (e.g., foundation models) that are pretrained on a massive data corpus and then adapted to solve a variety of downstream tasks. To summarize performance across multiple tasks, evaluation metrics are often aggregated into a summary metric, e.g., average accuracy across 10 question-answering tasks. When aggregating evaluation metrics, it is useful to incorporate uncertainty in the aggregate metric in order to gain a more realistic understanding of model performance. Our objective in this work is to demonstrate how statistical methodology can be used for quantifying uncertainty in metrics that have been aggregated across multiple tasks. The methods we emphasize are bootstrapping, Bayesian hierarchical (i.e., multilevel) modeling, and the visualization of task weightings that consider standard errors. These techniques reveal insights such as the dominance of a specific model for certain types of tasks despite an overall poor performance. We use a popular ML benchmark, the Visual Task Adaptation Benchmark (VTAB), to demonstrate the usefulness of our approaches.

性能评估不确定性量化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。