提出抗干扰评估框架,让模型性能估计算法在低方差下仍可靠。
Fault-Tolerant Evaluation for Sample-Efficient Model Performance Estimators
- 引入可调容差ε,同时考虑偏差与方差进行评估
- 自动优化ε值,使评估在不同方差环境下都稳定有效
- 适合需要低成本高可靠评估的AI服务部署场景
在模型即服务时代,组织依赖第三方AI模型快速部署。但新兴应用动态变化、新数据集持续涌现、模型性能宣称日益增多,导致高效可靠的模型服务验证愈发困难。这促使发展样本高效的性能估计算法,通过智能选择标注实例来降低标注成本。然而现有方法在低方差场景下表现不佳:均方误差(RMSE)混淆了偏差与方差,当方差小时会掩盖持续偏差;基于p值的检验则过于敏感,对微小偏离也拒绝合理估计器。为此,我们提出一种故障容错评估框架,将偏差与方差纳入可调容差ε的考量,实现对性能估计算法在实际可接受误差范围内的评估。理论证明ε的合理校准能确保在不同方差区间下评估可靠性,并进一步提出算法自动优化和选择ε。在真实数据集上的实验表明,该框架能提供全面且可操作的估计算法行为洞察。
原文摘要 · Abstract (English)
In the era of Model-as-a-Service, organizations increasingly rely on third-party AI models for rapid deployment. However, the dynamic nature of emerging AI applications, the continual introduction of new datasets, and the growing number of models claiming superior performance make efficient and reliable validation of model services increasingly challenging. This motivates the development of sample-efficient performance estimators, which aim to estimate model performance by strategically selecting instances for labeling, thereby reducing annotation cost. Yet existing evaluation approaches often fail in low-variance settings: RMSE conflates bias and variance, masking persistent bias when variance is small, while p-value based tests become hypersensitive, rejecting adequate estimators for negligible deviations. To address this, we propose a fault-tolerant evaluation framework that integrates bias and variance considerations within an adjustable tolerance level ${\varepsilon}$, enabling the evaluation of performance estimators within practically acceptable error margins. We theoretically show that proper calibration of ${\varepsilon}$ ensures reliable evaluation across different variance regimes, and we further propose an algorithm that automatically optimizes and selects ${\varepsilon}$. Experiments on real-world datasets demonstrate that our framework provides comprehensive and actionable insights into estimator behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。