用无标签学习量化并审计大模型评估的偏见,提升判断公正性。
Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

- 将人类评价中的正样本与未标注输出结合,构建偏见检测框架。
- 在多个数据集上显著降低冗长性偏差,更贴近人类偏好。
- 无需重训练,适合需要可解释评估结果的研究者使用。
大型语言模型(LLMs)被广泛用作可扩展评估的裁判,但这类‘大模型即裁判’系统存在与语义质量无关的系统性偏见,尤其是冗长性偏差。与此同时,人工监督成本高且通常只提供可靠正例,导致多数输出未被标注且质量混杂。本文将选择性人工监督下的大模型评估建模为正-未标记学习问题,提出基于部分最优传输的几何审计框架。通过在固定嵌入空间中对齐少量人工验证的正例与可靠的未标注输出,该方法识别出与人类一致的偏好,并在不重新训练的前提下纠正有偏的裁判。实验表明,该方法显著提升了与人类偏好的一致性,增强了对呈现方式偏差的鲁棒性,并提供可解释的置信度估计,为现有大模型评估流程提供了可扩展、统计严谨的替代方案。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Meanwhile, human supervision is costly and typically selective, yielding reliable positive judgments but leaving most outputs unlabelled and potentially mixed in quality. We formulate LLM evaluation under selective human supervision as a positive--unlabelled learning problem and propose a geometric auditing framework based on Partial Optimal Transport. By aligning a small set of human--verified positives with a reliable subset of unlabelled outputs in a fixed embedding space, our method identifies human--consistent preferences and corrects biased judges without retraining. Experiments demonstrate improved alignment with human preferences, increased robustness to presentation biases, and interpretable confidence estimates, offering a scalable and statistically grounded alternative to existing LLM--as--a--judge pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。