推理模型更准但有偏见,提出轻量方案缓解。
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
- 用显式评估计划引导模型判断,提升公正性
- 推理模型在复杂任务上准确率更高,抗攻击更强
- 适合需要公平评估的AI评测场景
本文首次系统比较大型推理模型(LRMs)与非推理大语言模型在判断任务中的表现。实验发现:1)在推理密集型任务中,LRMs的判断准确率显著优于非推理模型;2)LRMs具有更强的指令遵循能力;3)对针对判断任务的对抗攻击更具鲁棒性;4)但依然存在明显评估偏见。为此,我们提出PlanJudge——一种轻量级评估策略,通过让模型先生成显式评估计划再执行判断。实验表明,该方法能显著缓解偏见,同时保持整体判断准确率。
原文摘要 · Abstract (English)
This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. Our empirical analysis yields four key findings: 1) LRMs outperform non-reasoning LLMs in terms of judgment accuracy, particularly on reasoning-intensive tasks; 2) LRMs demonstrate superior evaluation instruction-following capabilities; 3) LRMs exhibit enhanced robustness against adversarial attacks targeting judgment tasks; 4) However, LRMs still exhibit strong evaluation biases. To mitigate this bias vulnerability, we propose PlanJudge, a lightweight evaluation strategy that prompts the model to generate an explicit evaluation plan before executing the judgment. Despite its simplicity, our experiments demonstrate that PlanJudge significantly mitigates biases in LLM-as-a-Judge while preserving overall judgment accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。