arXiv:2606.03043cs.CL2026-06

大模型评分共识不等于人类对齐,几何分析揭示其评分轴与人类严重偏离。

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

  • 通过几何量测发现模型评分轴与人类几乎正交,偏离度达87°以上。
  • 模型间评分一致性高(r≈0.35),但与人类一致性低(r≈0.27-0.32)。
  • 仅后验校准能提升对齐度,且仍不及人类自身可靠性。

当前大模型作为评分者已成标准,但模型间高度一致而与人类弱相关。我们在四个社区构建的印地语数据集、八种印地语语言及41个大模型上,测量了四种几何指标:评分分布范围、有效秩、与人类评分子空间的主角度、模型间与人类间的相关性,均带自举置信区间。在主观评价维度,模型评分范围不足人类的一半(σ_J / σ_H ≈ 0.3–0.5),其评分轴与人类几乎正交(87°–89°),远于人类间差异(78°–81°)。模型间一致性(r_LL ≈ 0.35)高于模型与人类一致性(r_LH ≈ 0.27–0.32)。在有可验证答案的事实类任务中,各项指标回归人类范围(主角58.5°;r_LH = 0.519)。微调和偏好优化虽恢复评分跨度(0.32 → 1.08),但评分轴未显著移动(仍为87°–88°)。唯有基于小规模人类锚定集的后验校准,可同时改善四项群体健康度量,使240亿参数的印地语模型(r=0.184)超越GPT-5.5(r=0.123),但仍低于人类自身可靠性(人类-人类r=0.474)。我们主张,模型间共识应仅在直接几何检查通过时视为人类对齐证据;否则反映的是退化子空间内的集体偏差。

原文摘要 · Abstract (English)

LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared signal or shared bias by measuring four geometric quantities on the standard LLM-as-judge stack across four community-built Indic datasets, eight Indic languages, and 41 LLM judges: score spread, effective rank, principal angle to the human subspace, and stacked correlations among judges and humans, all with bootstrap confidence intervals. On subjective rubrics, judges use less than half the human score range ($σ_J / σ_H \approx 0.3$--$0.5$). Their evaluation axis is nearly orthogonal to the human one and noticeably further from humans than humans are from each other ($87^\circ$--$89^\circ$ versus $78^\circ$--$81^\circ$). Inter-LLM agreement exceeds LLM--human agreement ($r_{LL} \approx 0.35$ versus $r_{LH} \approx 0.27$--$0.32$). On a rubric with a verifiable factual answer, the same diagnostics fall back into the human range (axis $58.5^\circ$; $r_{LH} = 0.519$). Fine-tuning and preference optimization recover spread ($0.32 \rightarrow 1.08$) but barely move the axis (still $87^\circ$--$88^\circ$). Only post-hoc calibration on a small human-anchored set improves all four community-health rubrics together, placing a calibrated 24B Indic judge ($r = 0.184$) ahead of GPT-5.5 ($r = 0.123$), yet still short of human reliability (human-human $r = 0.474$ on the verifiable rubric). We argue that inter-LLM agreement should be considered evidence of human alignment only when a direct geometric check on the judge's score subspace passes; otherwise, the consensus reflects agreement within a collapsed subspace.

大模型评测评分一致性几何对齐人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。