arXiv:2604.23099cs.LGcs.AI2026-04被引 1

用预训练高斯过程高效评估生成式AI,少用65倍样本仍更准且能发现更多问题。

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation

论文配图:ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
图 1 · 摘自论文原文
  • 用预训练高斯过程代理性能评分函数,实现快速估计。
  • 仅需8-65倍少样本即可达到99%准确率,且发现更多失败案例。
  • 适合需要低成本、高覆盖率评估的AI研发团队使用。

生成式AI评估因推理缓慢、评测成本高及模型与基准不断增多而日益资源密集。我们提出ProEval,一种主动评估框架,利用迁移学习高效估算性能并识别失败案例。ProEval采用预训练高斯过程(GPs)作为性能评分函数的代理,将模型输入映射为错误严重度或安全违规等指标。通过将性能估计建模为贝叶斯积分(BQ),将故障发现建模为超水平集采样,我们设计了具有不确定性的决策策略,主动选择或合成高度信息量的测试输入。理论上,我们证明了基于预训练GP的BQ估计器无偏且有界。在推理、安全对齐和分类基准上的大量实验表明,ProEval显著优于现有基线:只需8-65倍少样本即可实现与真实值相差1%以内的估计,同时在更严格的预算下揭示更多样化的失败案例。

原文摘要 · Abstract (English)

Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks. We propose ProEval, a proactive evaluation framework that leverages transfer learning to efficiently estimate performance and identify failure cases. ProEval employs pre-trained Gaussian Processes (GPs) as surrogates for the performance score function, mapping model inputs to metrics such as the severity of errors or safety violations. By framing performance estimation as Bayesian quadrature (BQ) and failure discovery as superlevel set sampling, we develop uncertainty-aware decision strategies that actively select or synthesize highly informative inputs for testing. Theoretically, we prove that our pre-trained GP-based BQ estimator is unbiased and bounded. Empirically, extensive experiments on reasoning, safety alignment, and classification benchmarks demonstrate that ProEval is significantly more efficient than competitive baselines. It requires 8-65x fewer samples to achieve estimates within 1% of the ground truth, while simultaneously revealing more diverse failure cases under a stricter evaluation budget.

生成式AI性能评估高斯过程主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。