通过配对评估揭示大模型判官在检索生成中的自偏见问题
Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG

- 构建答案配对矩阵,用同一模型作判官评估不同生成答案
- 发现同模型判官间召回率差异接近零,仅避免诱导结论的判断更准确
- 强调报告完整矩阵与标签一致性,避免误判虚假正例
LLM作为判官评估广泛用于源基检索增强生成(RAG),但使用相同模型族作为生成器和判官会导致自偏见难以识别。本文提出Eval-Pair Matrix,一种受控的元评估协议。基于GaRAGe问题与溯源文本,每条记录引入一个隐藏的答案因果矛盾,利用GPT、Grok和Gemini模型从扰动文本生成答案,并用相同模型作为盲评判官,评估答案与原始文本的一致性。实验包含300个核心记录、897个标注生成输出及2,683个判官判断,形成3×3交叉矩阵;主分析使用275个完全验证记录。不同于传统对比对角与非对角单元,本文通过配对同一候选答案的判官来估计同模型效应。结果表明:对角与非对角F1相似,同模型召回率差异接近零(-0.5个百分点;95%聚类自举置信区间[-2.7, +1.7])。唯一显著的配对差距是避免诱导主张的答案,其匹配判官标记率更低(-4.3个百分点)。人工评估显示,看似的假阳性实为替代源错误检测、诱导主张是否采纳的标注失误或模糊案例,无一被认定为真正的误报。方法论启示:RAG判官研究应报告完整矩阵、答案配对效应、行为分层及标签任务一致性。
原文摘要 · Abstract (English)
LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to identify. We introduce Eval-Pair Matrix, a controlled meta evaluation protocol for source-grounded RAG. Starting from GaRAGe questions and grounding passages, we induce one hidden answer-causal contradiction per record, generate answers from perturbed passages with GPT, Grok, and Gemini models, and then use the same models as blind judges to evaluate each answer against the original passages. The experiment contains 300 core records, 897 labeled generator outputs, and 2,683 judge verdicts in a crossed 3 x 3 matrix; the primary analysis uses 275 fully validated records. Instead of comparing diagonal and off-diagonal cells across different answers, we estimate same-model effects by pairing judges on the exact same candidate answer. This changes the interpretation: diagonal and off diagonal F1 are similar, and the paired same-model recall effect is near zero (-0.5 pp; 95% cluster bootstrap CI [-2.7, +1.7]). The only robust paired gap is lower matching-judge flagging for answers that avoided the induced claim (-4.3 pp). A targeted human evaluation finds that reviewed apparent false positives are alternate source-error detections, mistakes in labeling whether the induced claim was adopted, or unclear cases; none were adjudicated as genuine false alarms. The lesson is methodological: RAG judge studies should report full matrices, answer-paired effects, behavior strata, and label-task alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。