通过多预测联合优化,用更少预算实现更高精度的统计估计。
Multiple-Prediction-Powered Inference
- 融合多种数据源,动态分配资源以提升估计效率
- 在三大LLM评估场景中误差显著低于现有方法
- 适合需要低成本高精度评估的研究者使用
统计估计常面临高成本高质量测量与多种低成本代理测量之间的权衡。我们提出多预测驱动推断(MultiPPI):一种通过最优分配资源到多样数据源来构建统计高效估计的一般框架。该工作提供了关于MultiPPI估计器最小最大最优性、有限样本表现及渐近正态性的理论保证。在三个不同的大型语言模型评估场景中的实验表明,MultiPPI始终比现有基线获得更低的估计误差。这一优势源于其预算自适应的分配策略,能够通过学习模型间的复杂成本与相关性结构,智能组合部分模型。
原文摘要 · Abstract (English)
Statistical estimation often involves tradeoffs between expensive, high-quality measurements and a variety of lower-quality proxies. We introduce Multiple-Prediction-Powered Inference (MultiPPI): a general framework for constructing statistically efficient estimates by optimally allocating resources across these diverse data sources. This work provides theoretical guarantees about the minimax optimality, finite-sample performance, and asymptotic normality of the MultiPPI estimator. Through experiments across three diverse large language model (LLM) evaluation scenarios, we show that MultiPPI consistently achieves lower estimation error than existing baselines. This advantage stems from its budget-adaptive allocation strategy, which strategically combines subsets of models by learning their complex cost and correlation structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。