arXiv:2606.15127cs.LG2026-06被引 11

评测推理模型时,不仅要看答案对错,还要看是否承认偏见。

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

  • 设计双维度诊断:能否被偏见误导,是否主动指出偏见内容。
  • 相同条件下,GPT-4o承认偏见率仅13%,而Claude Sonnet高达75%。
  • 适用于教育、审计等需透明推理的负责任AI场景。

推理模型在教育工具、决策支持和审计流程中广泛应用,其推理过程本身也成为审查对象。传统仅以最终答案正确性评估模型存在盲区:两个模型可得相同答案分,却在是否识别并标注注入的偏见内容上差异显著。本文提出一种基于痕迹的最小化诊断方法,包含两个维度:敏感性(偏见是否导致原本正确的答案出错)与承认度(推理轨迹中是否出现预设标准的偏见表面提及)。在数千次带有偏见的GSM8K测试中,GPT-4o与Claude Sonnet~4的敏感性率分别为1.3%与1.2%,但承认率分别为13.0%与75.0%,表明即使表现相似,模型对偏见的认知能力仍存在巨大差异。

原文摘要 · Abstract (English)

Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students intermediate steps, decision-support systems may require human oversight, and audit workflows may inspect traces for misleading or biased input. In such settings, two responses can receive the same final-answer score while differing in whether the trace explicitly flags injected biasing content. Accuracy-only evaluation collapses these cases. We study this gap as a measurement blind spot for responsible evaluation and introduce a minimal trace-level diagnostic with two axes: \emph{susceptibility} (whether the bias breaks a previously correct answer) and \emph{acknowledgment} (whether the trace contains a rubric-defined surface reference to the injected content). Across thousands of biased GSM8K trials, GPT-4o and Claude Sonnet~4 have similar susceptibility rates ($1.3\%$ vs. $1.2\%$) but substantially different acknowledgment rates ($13.0\%$ vs. $75.0\%$) under the same rubric.

AI评估偏见检测推理链负责任AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。