arXiv:2604.02585cs.AIcs.CL2026-04

用新方法让大模型不被教师无关背景干扰,提升评估公正性。

Mitigating LLM biases toward spurious social contexts using direct preference optimization

  • 结合对比推理与偏好优化,改进DPO以减少无关背景影响。
  • 在7类虚假情境下,模型评分偏差最大降低1.48分(7分制),平均偏倚降84%。
  • 适合教育评估、高风险决策等需公平性的场景使用。

大语言模型在高风险决策中日益普及,但对虚假社会情境的敏感性可能引入有害偏见。本文以美国最大公开课堂转录数据集(NCTE)和专家评分为基础,评估七种前沿及开源模型在教师经验、学历、人口身份及奉承诱导框架等七类虚假上下文下的表现。结果发现,无关信息可使模型评分变动高达1.48分(7分制)。现有提示缓解策略及主流后训练方法(如SFT、DPO)均无法有效解决该问题。为此提出Debiasing-DPO,结合对比推理增强的DPO与专家标签SFT,既抑制虚假上下文影响又避免模式崩溃。在3-8B参数的Llama与Qwen Instruct模型上,该方法平均降低偏倚84%,预测准确率提升52%。研究揭示强模型可能更敏感,而Debiasing-DPO能同步提升准确率与鲁棒性。

原文摘要 · Abstract (English)

LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious context can introduce harmful biases. This is a critical concern when models are deployed for tasks like evaluating teachers' instructional quality, where biased assessment can affect teachers' professional development and career. We investigate model robustness to spurious social contexts about teachers using the largest publicly available dataset of U.S. classroom transcripts (NCTE) paired with expert evaluation scores. Evaluating seven frontier and open-weight models across seven categories of spurious contexts -- including teacher experience, education level, demographic identity, and sycophancy-inducing framings -- we find that irrelevant contexts can shift model-generated ratings by up to 1.48 points on a 7-point scale. Prompt-based mitigations and popular post-training methods, such as Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), fall short of addressing the issue. We propose Debiasing-DPO, which combines contrastive reasoning-augmented DPO with SFT on expert labels to reduce the effect of spurious context while avoiding mode collapse. Applied to Llama and Qwen Instruct models of 3-8B parameters, Debiasing-DPO reduces bias by 84% and improves predictive accuracy by 52% on average across models. Our findings from the educational dataset highlight that stronger models may exhibit greater sensitivity despite higher accuracy, and Debiasing-DPO can improve both accuracy and robustness in prompt-based prediction tasks.

大模型偏见偏好优化教育评估鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。