arXiv:2604.05460stat.MEcs.AI2026-04被引 1

用张量补全方法为大模型评估提供更可靠的不确定性量化。

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency

  • 将大模型评估建模为低秩张量的半参数推断问题。
  • 提出一种去偏估计器,实现渐近正态性与最优样本复杂度。
  • 适合关注评估可靠性、统计推断的算法研究者和评测平台设计者。

大型语言模型(LLM)评估平台日益依赖成对的人类判断。这些数据噪声大、稀疏且分布不均,但排行榜报告中缺乏足够的不确定性量化。本文将其建模为在Bradley-Terry-Luce型模型下,通过成对比较观测到的低秩潜在得分张量的半参数推断问题。该设定引入了结构化观测、非均匀采样和成对对比的新张量补全框架。目标是估计光滑函数ψ(T⋆),包括线性估计量(如能力差值)和非线性估计量(如胜率)。我们推导出低秩切空间上的信息算子、高效影响函数及半参数效率界,并构造出具有渐近正态性的一步去偏估计器。核心挑战在于信息算子各向异性且不与切空间投影交换,导致传统方法失效。为此,我们提出得分白化方法,均衡局部Fisher信息,恢复最优样本复杂度下的稳定推断。结果为大模型评估中的不确定性量化提供了理论基础,也为从成对数据中推断低秩结构提供了通用框架。

原文摘要 · Abstract (English)

Large language model (LLM) evaluation platforms increasingly rely on pairwise human judgments. These data are noisy, sparse, and non-uniform, yet leaderboards are reported with limited uncertainty quantification. We study this as semiparametric inference for a low-rank latent score tensor observed through pairwise comparisons under Bradley-Terry-Luce-type models. This places LLM evaluation in a new tensor completion setting with structured observations, non-uniform sampling, and pairwise contrasts. Our target is a smooth functional $ψ(T^\star)$, including linear estimands such as ability gaps and nonlinear ones such as win probabilities. We derive the information operator on the low-rank tangent space, the efficient influence function, and the semiparametric efficiency bound, then construct a one-step debiased estimator with asymptotic normality. A central challenge is that the information operator is anisotropic and does not commute with the tangent-space projection, creating a bottleneck absent from isotropic models. We introduce a score-whitening method that equalizes local Fisher information and restores stable inference at the optimal sample-complexity scale. Our results provide a principled framework for uncertainty quantification in LLM evaluation and more broadly for inference on low-rank structures from pairwise data.

大模型评估张量补全统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。