arXiv:2608.16210cs.LGstat.ML2026-08

用廉价信号提升大模型条件评估精度,省时省力还更准。

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

  • 通过局部中心化处理廉价信号,消除偏差并提升估计效率。
  • 在多个数据集上显著降低评估误差,最高提升达37%。
  • 适合需要高精度条件评估的AI研究者与工程团队使用。

整体准确率掩盖了模型在不同情况下的表现差异。仅用黄金标签估算条件性能成本高昂,而大语言模型评分、置信度等廉价辅助信号可批量获取,但常存在偏差或校准不准问题。本文提出LACE(局部增强控制变量评估)方法,核心是局部中心化:在目标条件区域内减去廉价信号的条件均值后,任意线性增强项的条件均值为零,不影响估计量。通过结合标注子集的黄金标签残差均值和全样本廉价信号均值,构造局部岭形控制变量,实现无需校准的识别、分组条件下无偏估计及局部最优效率。效率增益由总体局部$R^2$决定,反映不同条件值下廉价信号的增益潜力。同时推导出直接成对模型差距与部署加权分数的估计器。在MATH-500、ScienceQA、MMLU、WinoGrande、HellaSwag、TruthfulQA、GSM8K和ARC共8个数据集上进行了实证评估。

原文摘要 · Abstract (English)

Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. We propose LACE (Local Augmented Control-Variate Evaluation), a semi-supervised estimator for conditional LLM evaluation. The key step is local centering: after subtracting the conditional mean of a cheap signal within the target profile region, any linear augmentation has zero conditional mean and therefore cannot change the estimand. The augmentation coefficient is used only for efficiency, and a local ridge control variate combines a gold-label residual mean from the labeled subset with a cheap-signal mean from the full item pool. We prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient. The resulting gain formula is governed by a population local $R^2$, which characterizes how the efficiency attainable from the cheap signals varies across profile values. We also derive corresponding estimators for direct paired model gaps and deployment-weighted scores. We empirically evaluate the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC.

模型评估控制变量条件分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。