提出统一框架,让大模型评价更贴近人类判断。
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- 建模人类偏好为隐变量,线性拟合大模型评分偏差
- 在两个数据集上提升与人类评分的一致性,降低差距
- 适合评估大模型生成质量的研究者和开发者
大语言模型作为评判者(LLM-as-a-judge)被广泛用于大规模评估模型输出,但其评分常与人类判断系统性偏离。本文提出Bridge,一个统一的统计框架,可在绝对评分与成对比较两种范式下显式桥接人类与大模型评价。该框架假设每个提示-响应对存在潜在的人类偏好分,并将大模型偏差建模为协变量的线性变换,捕捉差异来源。这提供了一个简洁且原理清晰的框架,用于修正大模型评分并刻画人机评价间的系统性差异。我们设计了高效拟合算法,并具备渐近统计推断保证。在六个大模型评判者和两个基准数据集(BigGen Bench与Chatbot Arena)上,Bridge显著提升了与人类评分的一致性(准确率、校准度、KL散度),并揭示了系统性人机差距。
原文摘要 · Abstract (English)
Large language models are increasingly used as judges (LLM-as-a-judge) to evaluate model outputs at scale, but their assessments often diverge systematically from human judgments. We present Bridge, a unified statistical framework that explicitly bridges human and LLM evaluations under both absolute scoring and pairwise comparison paradigms. Bridge posits a latent human preference score for each prompt-response pair and models LLM deviations as linear transformations of covariates that capture sources of discrepancies. This offers a simple and principled framework for refining LLM ratings and characterizing systematic discrepancies between humans and LLMs. We provide an efficient fitting algorithm with asymptotic guarantees for statistical inference. Using six LLM judges and two benchmarks (BigGen Bench and Chatbot Arena), Bridge achieves higher agreement with human ratings (accuracy, calibration, and KL divergence) and exposes systematic human-LLM gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。