用身份映射提升无标签大模型评估的准确性与效率
Identity-Link IRT for Label-Free LLM Evaluation: Preserving Additivity in TVD-MI Scores
- 采用恒等映射替代传统非线性链接,保持评分可加性
- 在33%评估覆盖率下,误差仅0.117,排名相关性高达0.972
- 适用于大模型评估及其他有界响应场景,对评委鲁棒
使用总变差距离互信息(TVD-MI)进行大语言模型成对比较时,每次试验产生二值评判结果。我们发现,对这些二值结果取均值可得具有可加结构的中心化概率评分,适合直接应用于项目反应理论(IRT),无需非线性链接函数。传统最大似然IRT方法使用逻辑链接,但实证显示其引入曲率破坏可加性:在三个领域中,恒等链接在原始数据上的中位曲率仅为0.080-0.150(第95百分位[0.474, 0.580]),而probit/logit则显著更高(中位[0.245, 0.588],第95百分位[0.825, 2.252])。该截断线性模型源于吉尼熵最大化,构建了盒约束最小二乘形式以处理边界饱和。在33%覆盖下,保持了0.117±0.008的留出集均方根误差,且保留了代理排序(斯皮尔曼ρ=0.972±0.015),评估次数仅为完整密集评估的三分之一。法官鲁棒性分析(GPT-4o-mini vs. Llama3-70b)显示代理排序高度一致(ρ=0.872),并验证恒等链接优势。TVD-MI的几何结构通过恒等映射最佳保留,实现高效的大模型评估,适用于其他有界响应领域。
原文摘要 · Abstract (English)
Pairwise comparisons of large language models using total variation distance mutual information (TVD-MI) produce binary critic decisions per pair. We show that averaging TVD-MI's binary trials yields centered-probability scores with additive structure suitable for item-response theory (IRT) without nonlinear link functions. Maximum-likelihood approaches to IRT use logistic links, but we find empirically that these transformations introduce curvature that breaks additivity: across three domains, the identity link yields median curl on raw data of 0.080-0.150 (P95 = [0.474, 0.580]), whereas probit/logit introduce substantially higher violations (median [0.245, 0.588], P95 [0.825, 2.252]). We derive this clipped-linear model from Gini entropy maximization, yielding a box-constrained least-squares formulation that handles boundary saturation. At 33% coverage, we achieve holdout RMSE $0.117 \pm 0.008$ while preserving agent rankings (Spearman $ρ= 0.972 \pm 0.015$), three times fewer evaluations than full dense. Judge robustness analysis (GPT-4o-mini vs. Llama3-70b) shows strong agreement in agent rankings ($ρ= 0.872$) and consistent identity-link advantage. TVD-MI's geometry is best preserved by identity mapping for efficient LLM evaluation, applicable to other bounded-response domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。