arXiv:2507.02799cs.CL2025-07被引 2

推理模型反而更易被诱导出偏见,挑战了‘推理=更安全’的假设。

Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models

  • 用基准测试对比推理模型与基础模型的偏见脆弱性。
  • 显式推理模型比提示推理模型更易被攻击,尤其在故事类提示下。
  • 揭示推理机制可能无意中强化刻板印象,需改进设计。

推理语言模型(RLMs)通过链式思维(CoT)提示或微调推理轨迹,具备处理复杂多步任务的能力。尽管这提升了可靠性,但其对社会偏见的鲁棒性仍不明确。本文利用原为大语言模型设计的CLEAR-Bias基准,系统评估前沿RLMs在多元社会文化维度上的抗偏见能力。采用LLM作为裁判进行自动化安全评分,并运用越狱技术测试内置安全机制强度。研究聚焦三个问题:(i)推理能力如何影响模型公平性与鲁棒性;(ii)微调推理模型是否比依赖推理提示的模型更安全;(iii)不同推理机制下越狱攻击的成功率差异。结果表明,显式推理模型普遍比无推理机制的基础模型更易受偏见诱导,暗示推理可能无意中打开刻板印象强化的新路径。微调推理模型表现优于仅靠推理提示的模型,后者尤其易受叙事提示、虚构人格或奖励引导指令的上下文重构攻击。研究挑战了‘推理提升鲁棒性’的假设,强调需发展更具偏见意识的推理设计。

原文摘要 · Abstract (English)

Reasoning Language Models (RLMs) have gained traction for their ability to perform complex, multi-step reasoning tasks through mechanisms such as Chain-of-Thought (CoT) prompting or fine-tuned reasoning traces. While these capabilities promise improved reliability, their impact on robustness to social biases remains unclear. In this work, we leverage the CLEAR-Bias benchmark, originally designed for Large Language Models (LLMs), to investigate the adversarial robustness of RLMs to bias elicitation. We systematically evaluate state-of-the-art RLMs across diverse sociocultural dimensions, using an LLM-as-a-judge approach for automated safety scoring and leveraging jailbreak techniques to assess the strength of built-in safety mechanisms. Our evaluation addresses three key questions: (i) how the introduction of reasoning capabilities affects model fairness and robustness; (ii) whether models fine-tuned for reasoning exhibit greater safety than those relying on CoT prompting at inference time; and (iii) how the success rate of jailbreak attacks targeting bias elicitation varies with the reasoning mechanisms employed. Our findings reveal a nuanced relationship between reasoning capabilities and bias safety. Surprisingly, models with explicit reasoning, whether via CoT prompting or fine-tuned reasoning traces, are generally more vulnerable to bias elicitation than base models without such mechanisms, suggesting reasoning may unintentionally open new pathways for stereotype reinforcement. Reasoning-enabled models appear somewhat safer than those relying on CoT prompting, which are particularly prone to contextual reframing attacks through storytelling prompts, fictional personas, or reward-shaped instructions. These results challenge the assumption that reasoning inherently improves robustness and underscore the need for more bias-aware approaches to reasoning design.

偏见检测推理模型安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。