揭示大模型评分偏见的隐藏机制,用激活空间几何解释并控制评分偏差。
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

- 从隐藏层激活态分析评分偏见,发现其存在于低维特定子空间中。
- 沿该子空间调控可逆转或重现评分偏差,效果比随机方向强十倍。
- 仅用线性投影即可预测模型在新任务上的失败,适合优化评分系统。
现有研究多从输入输出层面分析大模型作为评判者的评分偏见:扰动输入、测量分数变化并提出提示缓解方案。本文主张,同样的偏见可在评判模型的隐藏状态层面以表征形式呈现,与输入输出视角互补且更具操作价值。我们在七种模型、七类偏见和九个基准上报告三项发现:几何层面,基线输入激活集中在紧凑流形,而有偏输入则沿低维、类型特异的子空间偏移,且随网络深度增强,三种估计器均能稳定恢复该子空间;因果层面,沿该子空间操控隐藏状态可双向驱动评分,正向偏移再现有偏评分,反向偏移还原基线评分,而等范数随机方向仅产生约十分之一的效果;操作层面,仅通过简单线性投影到同一偏见方向特征,即可准确预测模型在三个全新基准上的评分失败,显著优于基于文本的替代方法。将偏见视为激活空间的几何结构,统一了几何形态、因果控制与预测能力,形成完整框架。
原文摘要 · Abstract (English)
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven judges, seven bias types, and nine benchmarks. Geometry: baseline judging inputs occupy a tight activation manifold while biased inputs are displaced along a low-dimensional, type-specific subspace that sharpens with depth and is recovered consistently by three families of estimators. Causal control: steering hidden states along this subspace drives scoring in both directions, forward shifts reproducing biased scoring on clean inputs and reverse shifts restoring baseline scoring on biased ones, while matched-norm random directions produce shifts an order of magnitude smaller. Operational: a simple linear projection onto the same bias-direction features anticipates judge failures on three entirely unseen benchmarks, substantially outperforming text-based alternatives. Reading bias as activation geometry, rather than as input-output noise, unifies geometric structure, causal control, and operational prediction within a single framework. The project page is available at https://xzx34.github.io/unfair-judge/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。