用自适应方法融合真实与合成数据,提升大模型评估的可靠性与效率。
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
- 动态调节合成数据权重,自动切换至传统评估方式以保证精度。
- 在量化、提示设计等任务中,样本效率优于或至少不劣于传统方法。
- 适合需要高可靠且高效评估的大模型优化场景。
从多个候选人工智能模型(如大语言模型)中选择最优模型需准确的性能估计。理想情况下,这依赖大量真实世界数据的实证评估,但此类评估成本高昂且难以规模化。为应对挑战,自动评估方法利用由自动化评估器(如大语言模型作为评判者)生成的合成数据,降低了方差但可能引入偏差。近期方法采用半监督预测驱动推断(PPI)来纠正自动评估器的偏差。然而,实际使用中,自动评估器可能导致样本效率低于仅使用真实数据的传统方法。本文提出R-AutoEval+框架,在模型评估中提供有限样本下的可靠性保障,同时确保样本效率优于或至少不劣于传统方法。其核心创新在于自适应构建评估变量,动态调整对合成数据的依赖程度,当自动评估器准确性不足时自动回退至传统方法。在使用大语言模型作为评判者优化模型权重量化设置、提示设计及推理预算分配的任务上,实验验证了R-AutoEval+在可靠性与效率上的优越性。
原文摘要 · Abstract (English)
Selecting artificial intelligence (AI) models, such as large language models (LLMs), from multiple candidates requires accurate performance estimation. This is ideally achieved through empirical evaluations involving abundant real-world data. However, such evaluations are costly and impractical at scale. To address this challenge, autoevaluation methods leverage synthetic data produced by automated evaluators, such as LLMs-as-judges, reducing variance but potentially introducing bias. Recent approaches have employed semi-supervised prediction-powered inference (PPI) to correct for the bias of autoevaluators. However, the use of autoevaluators may lead in practice to a degradation in sample efficiency compared to conventional methods using only real-world data. In this paper, we propose R-AutoEval+, a novel framework that provides finite-sample reliability guarantees on the model evaluation, while also ensuring an enhanced (or at least no worse) sample efficiency compared to conventional methods. The key innovation of R-AutoEval+ is an adaptive construction of the model evaluation variable, which dynamically tunes its reliance on synthetic data, reverting to conventional methods when the autoevaluator is insufficiently accurate. Experiments on the use of LLMs-as-judges for the optimization of quantization settings for the weights of an LLM, for prompt design in LLMs, and for test-time reasoning budget allocation in LLMs confirm the reliability and efficiency of R-AutoEval+.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。