多语言大模型评价存在排名颠倒问题,本文提出无标签校准方法解决此难题。
Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator

- 通过双中心化处理评分矩阵,分离语言与模型的交互影响。
- 在7920次测试中,排名一致性从0.650提升至0.902,显著增强稳定性。
- 无需人工标注即可校准,适合多语言评测场景的公平性分析。
多语言大模型评判者在不同语言提示下产生不一致的模型排名:在涵盖八种语言的Agent-as-a-Judge基准上,顶级模型随语言交替变化,15组模型对中有7组呈现统计显著的排名反转。我们将此视为测量偏差问题,发现多语言评分可分解为任务难度、模型能力与语言-模型交互项三部分,其中交互项可通过无监督双中心化恢复。我们提出该估计器(共识校准,CBC),给出$O(1/ ext{sqrt}{n})$的有限样本浓度界,方差常数为$(1- frac{1}{m})(1- frac{1}{k})$,并证明其在任务-语言交互存在时仍无偏。在7,920次评测(6个模型、8种语言、55个任务、3种框架)中,CBC将跨任务排名一致性$τ$从0.650提升至0.902,且每语言决策与保留集加性模型最优解完全一致(100%对比原始仅68.5%)。在独立收集的M-RewardBench数据集(7语言、每语言1,500项,共10,500实例、5个评估者)中,与公开人类黄金偏好的一致性从68.7%提升至76.6%(提升7.9个百分点,95%置信区间[6.0, 9.9]),为下游效用最强外部证据。该方法本质是带零和约束的双因素方差分析交互项恢复,贡献在于将其作为无标签后处理校准器应用于多语言大模型评判,提供显式有限样本界与鲁棒无偏性保证。
原文摘要 · Abstract (English)
Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Turkish, and Swahili, and 7 of 15 backbone pairs show statistically significant pairwise rank reversal. We treat this as a measurement problem. The multilingual judge score decomposes additively into task difficulty, backbone skill, and a language-backbone interaction term, the last of which is recoverable without human labels by double-centering the cell-mean score matrix. We make this estimator (\textbf{Consensus-Based Calibration}, CBC) explicit, give an $O(1/\sqrt{n})$ finite-sample concentration bound with variance constant $(1-\tfrac{1}{m})(1-\tfrac{1}{k})$, and show that it is unbiased even when task-language interactions are present. Across 7{,}920 judge runs (6 backbones, 8 languages, 55 tasks, 3 frameworks), CBC raises held-out cross-task rank consistency $τ$ from 0.650 to 0.902 and agrees with the held-out additive-model oracle in 100\% of per-language decisions versus 68.5\% raw; these are consistency diagnostics, not human-grounded correctness measures. On a separately collected M-RewardBench panel (7 languages, 1{,}500 items per language, 10{,}500 language-item instances, 5 evaluators), panel agreement with the public human gold preferences rises from 68.7\% to 76.6\% (gain 7.9 percentage points, 95\% CI $[6.0, 9.9]$), our strongest external evidence of downstream usefulness. The estimator is the standard two-way ANOVA interaction-recovery operation under sum-to-zero contrasts; our contribution is its application as a label-free post-hoc calibrator for multilingual LLM judges, an explicit finite-sample concentration bound, and an unbiasedness result that holds even under task-language misspecification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。