arXiv:2505.12301cs.AIcs.CL2025-05被引 4

让大模型评价结果更贴近人类真实判断分布,提升评测可靠性。

Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge

  • 用分布对齐替代单点打分,更准确捕捉人类评价的多样性。
  • 在多个任务上超越闭源大模型和传统方法,提升评测准确率与鲁棒性。
  • 适合需要高可信度自动评估的研究者和开发者使用。

大模型作为评判者在大模型评判范式中展现出强大能力,相比人工评判更具效率和灵活性。然而,现有方法主要依赖单点评价,忽略了人类评判中的固有差异与不确定性,导致信息损失,降低评估可靠性。为此,我们提出一种新型训练框架,显式对齐大模型生成的评判分布与真实人类分布。具体地,采用基于KL散度的分布对齐目标,并引入辅助交叉熵正则化以稳定训练过程。考虑到真实分布可能来自有限的人工标注,我们还引入对抗训练,增强模型对分布扰动的鲁棒性。在多种大模型主干架构和评估任务上的大量实验表明,该框架显著优于现有闭源大模型及传统单点对齐方法,具备更高的分布对齐质量、评估准确性和鲁棒性。

原文摘要 · Abstract (English)

LLMs have emerged as powerful evaluators in the LLM-as-a-Judge paradigm, offering significant efficiency and flexibility compared to human judgments. However, previous methods primarily rely on single-point evaluations, overlooking the inherent diversity and uncertainty in human evaluations. This approach leads to information loss and decreases the reliability of evaluations. To address this limitation, we propose a novel training framework that explicitly aligns the LLM-generated judgment distribution with empirical human distributions. Specifically, we propose a distributional alignment objective based on KL divergence, combined with an auxiliary cross-entropy regularization to stabilize the training process. Furthermore, considering that empirical distributions may derive from limited human annotations, we incorporate adversarial training to enhance model robustness against distribution perturbations. Extensive experiments across various LLM backbones and evaluation tasks demonstrate that our framework significantly outperforms existing closed-source LLMs and conventional single-point alignment methods, with improved alignment quality, evaluation accuracy, and robustness.

大模型评测分布对齐自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。