提出自动去偏的上下文偏好推理方法,提升大模型评估准确性。
Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation
- 基于费雪随机游走设计加权残差平衡,实现自动去偏估计。
- 在多种上下文场景下评估大语言模型,结果更准确高效。
- 适用于灵活深度学习方法,适合模型评估与分布外推场景。
为满足大语言模型严格且可扩展的评估需求,本文研究跨领域上下文依赖偏好得分函数的成对比较推理问题。聚焦上下文布拉德利-特里-卢斯模型,提出一种半参数高效估计器,通过聚合比较图中加权残差平衡项实现自动去偏估计。当权重由新颖的费雪随机游走策略导出时,可达到效率最优。进一步提出一种计算可行的方法,通过扰动权重函数的潜在表示计算权重。所提推断过程适用于一般得分函数估计器,兼容实践者使用的灵活深度学习方法。扩展至多重假设检验,采用高斯乘子自举法控制族错误率;针对分布偏移,引入交叉拟合重要性采样调整实现目标域推断。数值实验,包括多种上下文下的语言模型评估,验证了方法的准确性、效率与实用性。
原文摘要 · Abstract (English)
Motivated by the need for rigorous and scalable evaluation of large language models, we study contextual preference inference for pairwise comparison functionals of context-dependent preference score functions across domains. Focusing on the contextual Bradley-Terry-Luce model, we develop a semiparametric efficient estimator that automates the debiased estimation through aggregating weighted residual balancing terms across the comparison graph. We show that the efficiency is achieved when the weights are derived from a novel strategy called Fisher random walk. We also propose a computationally feasible method to compute the weights by a potential representation of nuisance weight functions. We show our inference procedure is valid for general score function estimators accommodating the practitioners' need to implement flexible deep learning methods. We extend the procedure to multiple hypothesis testing using a Gaussian multiplier bootstrap that controls familywise error and to distributional shift via a cross-fitted importance-sampling adjustment for target-domain inference. Numerical studies, including language model evaluations under diverse contexts, corroborate the accuracy, efficiency, and practical utility of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。