arXiv:2602.06625cs.CL2026-02被引 3

让大模型当裁判更公平可靠,自动规避位置、长度等干扰因素

FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

  • 把评分行为建模为可学习的策略,动态适应不同任务需求
  • 在多个基准上提升判断一致性与准确率,减少非语义偏差
  • 适合需要公正评估生成内容的研究者与评测系统开发者

现有大模型作为裁判的系统存在三大缺陷:难以适应任务和领域特定评价标准、受位置、长度、格式等非语义线索影响产生系统性偏差、在不同评估模式下表现不一致。为此,我们提出FairJudge,一种自适应、去偏且一致的大模型裁判框架。不同于将裁判视为静态评估者,FairJudge将评判行为建模为可学习且正则化的策略。从数据角度,构建高信息密度的裁判数据集,显式注入与评价行为对齐的监督信号。基于该数据集,采用课程式SFT-DPO-GRPO训练范式,逐步对齐评分准则遵循性、偏差缓解与跨模式一致性,同时避免灾难性遗忘。在多个内部及公开基准上的实验表明,FairJudge在一致性与F1指标上均显著优于基线,有效降低非语义偏差,并超越更大规模的指令微调大模型。所有资源将在论文接收后公开发布,以促进后续研究。

原文摘要 · Abstract (English)

Existing LLM-as-a-Judge systems suffer from three fundamental limitations: limited adaptivity to task- and domain-specific evaluation criteria, systematic biases driven by non-semantic cues such as position, length, format, and model provenance, and evaluation inconsistency that leads to contradictory judgments across different evaluation modes (e.g., pointwise versus pairwise). To address these issues, we propose FairJudge, an adaptive, debiased, and consistent LLM-as-a-Judge. Unlike prior approaches that treat the judge as a static evaluator, FairJudge models judging behavior itself as a learnable and regularized policy. From a data-centric perspective, we construct a high-information-density judging dataset that explicitly injects supervision signals aligned with evaluation behavior. Building on this dataset, we adopt a curriculum-style SFT-DPO-GRPO training paradigm that progressively aligns rubric adherence, bias mitigation, and cross-mode consistency, while avoiding catastrophic forgetting. Experimental results on multiple internal and public benchmarks show that FairJudge consistently improves agreement and F1, reduces non-semantic biases, and outperforms substantially larger instruction-tuned LLMs. All resources will be publicly released after acceptance to facilitate future research.

大模型评测去偏一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。