揭示大模型评判者在解释时对无关线索的依赖,提出改进方法提升判断公正性。
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

- 设计多种线索干扰实验,检验模型评分与解释是否受无关信息影响
- 发现模型在1000份摘要评估中存在显著解释漂移,尤其在误导性线索下
- 提出证据先行策略,有效减少模型对冗余和自信线索的依赖
大型语言模型越来越多地被用作摘要和对话评估的自动评判者。已有研究指出位置、冗长度和风格偏好等偏差,但多关注结果,忽视评判者的解释过程。本文探究大模型评判者是否具备线索不变性——即在保持原文不变的前提下,当非证据性线索被扰动时,其评分与解释是否保持稳定。我们引入盲测、真相、翻转、安慰剂、揭示后等五类线索干预手段,以及绑定意识的度量指标,量化结果锚定与理由锚定,包括标签对齐修辞、解释漂移,并进行一致性与刻板印象侵入检查。通过设计基于冗长度和置信度线索的锚定攻击,对比结构化思维链提示与证据锁-评分-排序(PROOF-BEFORE-PREFERENCE)两种缓解策略。基于1000条传统抽取式模型与LLM生成的摘要新数据集,发现标签与安慰剂扰动下存在显著线索锚定现象,而PROOF-BEFORE-PREFERENCE显著优于基线,提升线索不变性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such as position, verbosity, and style preferences, but largely focuses on outcomes, leaving judge explanations underexplored. We instead ask whether LLM judges are cue-invariant, i.e., whether their rankings and explanations remain stable when non-evidential cues are perturbed while holding the underlying texts fixed. We introduce a suite of cue interventions (Blind, Truth, Flip, Placebo, Reveal-After) and tie-aware metrics that quantify outcome anchoring and rationale anchoring, including label-aligned rhetoric and explanation drift, alongside consistency and stereotype-intrusion checks. We design anchoring attacks using verbosity and confidence cues, and compare two mitigations: structured chain-of-thought prompting and PROOF-BEFORE-PREFERENCE (evidence lock, score, rank). Using a new dataset of 1,000 summaries from traditional extractive models and LLMs, we find substantial cue-anchored rationalization under label and placebo perturbations, while PROOF-BEFORE-PREFERENCE markedly improves cue invariance over baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。