用智能分配预算的方法,低成本高效评估生成式AI质量。
Cost-Optimal Active AI Model Evaluation
- 根据成本与准确性动态分配弱评器和强评器的标注任务。
- 在相同精度下,总标注成本降低显著,尤其在样本难度差异大时。
- 适合需要频繁迭代、预算受限的AI模型评估场景。
生成式AI系统的开发周期中,持续评估、数据获取与标注耗费大量资源与时间。实践中,为快速迭代常依赖低成本的合成标注数据,但可能引入显著偏差。本文提出新型成本感知方法,主动平衡廉价但不准确的弱评器(如自动评分模型)与昂贵但更准确的强评器(如人工评估)的使用。目标是在给定标注预算下,以最小方差无偏估计强评器的均值。基于主动学习与预测驱动统计推断的最新研究,我们推导出一族成本最优策略,用于在弱评器与强评器间分配预算以最大化统计效率。通过合成与真实数据实验,我们验证了该策略在高难度变异性任务中可实现相同估计精度而大幅降低总预算,优于传统方法。
原文摘要 · Abstract (English)
The development lifecycle of generative AI systems requires continual evaluation, data acquisition, and annotation, which is costly in both resources and time. In practice, rapid iteration often makes it necessary to rely on synthetic annotation data because of the low cost, despite the potential for substantial bias. In this paper, we develop novel, cost-aware methods for actively balancing the use of a cheap, but often inaccurate, weak rater -- such as a model-based autorater that is designed to automatically assess the quality of generated content -- with a more expensive, but also more accurate, strong rater alternative such as a human. More specifically, the goal of our approach is to produce a low variance, unbiased estimate of the mean of the target "strong" rating, subject to some total annotation budget. Building on recent work in active and prediction-powered statistical inference, we derive a family of cost-optimal policies for allocating a given annotation budget between weak and strong raters so as to maximize statistical efficiency. Using synthetic and real-world data, we empirically characterize the conditions under which these policies yield improvements over prior methods. We find that, especially in tasks where there is high variability in the difficulty of examples, our policies can achieve the same estimation precision at a far lower total annotation budget than standard evaluation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。