arXiv:2504.09946cs.CYcs.CL2025-04被引 38

发现推理模型在判别任务中仍存偏见,提出有效缓解策略。

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

  • 构建基准测试,系统评估大模型判别偏见。
  • 推理模型对事实类任务更鲁棒,但存在位置与表面反思偏见。
  • 自省机制显著降低偏见,适合部署为自动评判系统。

大型推理模型(LRMs)如DeepSeek-R1和OpenAI-o1展现出卓越的推理能力,引发了其在大语言模型作为裁判(LLM-as-a-judge)场景中是否存在偏见的重要问题。本文构建全面基准,比较了大语言模型(LLMs)与推理模型在主观偏好对齐数据集和客观事实数据集上的判别偏见。通过分析从众、权威、位置和干扰偏见,揭示四大发现:(1) 尽管具备先进推理能力,LRMs仍易受上述偏见影响;(2) LRMs在事实相关数据集上表现出优于LLMs的鲁棒性;(3) LRMs存在明显位置偏见,偏好靠后选项;(4) 发现一种新偏见——“表面反思偏见”,即模仿推理的短语(如“等等,让我想想……”)显著影响模型判断。为此,设计并评估三种缓解策略:专用系统提示可使偏好对齐数据集偏见降低最多19%,事实数据集降低14%;上下文学习在偏好任务中提升最高达27%,但在事实任务中表现不一;自省机制在偏好数据集减少偏见最多10%,事实数据集减少16%,且对LRMs尤为有效。本研究为构建更可靠的自动判别框架提供关键洞见,尤其在LRMs日益被用作自动裁判的背景下。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) like DeepSeek-R1 and OpenAI-o1 have demonstrated remarkable reasoning capabilities, raising important questions about their biases in LLM-as-a-judge settings. We present a comprehensive benchmark comparing judging biases between LLMs and LRMs across both subjective preference-alignment datasets and objective fact-based datasets. Through investigation of bandwagon, authority, position, and distraction biases, we uncover four key findings: (1) despite their advanced reasoning capabilities, LRMs remain susceptible to the above biases; (2) LRMs demonstrate better robustness than LLMs specifically on fact-related datasets; (3) LRMs exhibit notable position bias, preferring options in later positions; and (4) we identify a novel "superficial reflection bias" where phrases mimicking reasoning (e.g., "wait, let me think...") significantly influence model judgments. To address these biases, we design and evaluate three mitigation strategies: specialized system prompts that reduce judging biases by up to 19\% in preference alignment datasets and 14\% in fact-related datasets, in-context learning that provides up to 27\% improvement on preference tasks but shows inconsistent results on factual tasks, and a self-reflection mechanism that reduces biases by up to 10\% in preference datasets and 16\% in fact-related datasets, with self-reflection proving particularly effective for LRMs. Our work provides crucial insights for developing more reliable LLM-as-a-Judge frameworks, especially as LRMs become increasingly deployed as automated judges.

模型偏见自动评判推理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。