arXiv:2512.23430cs.CL2025-12ACL被引 6

提出C2PO框架,统一解决大模型的刻板与结构化偏见

C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs

  • 通过因果反事实信号识别输入中的虚假关联特征
  • 在优化过程中同时抑制偏见特征并保持推理能力
  • 适合关注公平性与模型可靠性研究者使用

大语言模型中的偏见威胁可信度,主要表现为刻板印象(如性别、种族)和结构性偏见(如词汇重叠、位置偏好)。现有方法常孤立处理,往往治一伤一。本文系统分析发现,根源在于输入中潜藏的虚假特征相关性,诱发错误推理捷径。为此,提出因果-对比偏好优化(C2PO)框架,通过因果反事实信号分离偏见诱导特征,并采用公平敏感的偏好更新机制,在逻辑层动态评估贡献并抑制捷径特征。在涵盖刻板偏见(BBQ、Unqover)、结构性偏见(MNLI、HANS、Chatbot、MT-Bench)、跨域公平性(StereoSet、WinoBias)及通用能力(MMLU、GSM8K)的多基准测试中,C2PO有效缓解两类偏见,同时保持强推理性能。

原文摘要 · Abstract (English)

Bias in Large Language Models (LLMs) poses significant risks to trustworthiness, manifesting primarily as stereotypical biases (e.g., gender or racial stereotypes) and structural biases (e.g., lexical overlap or position preferences). However, prior paradigms typically address these in isolation, often mitigating one at the expense of exacerbating the other. To address this, we conduct a systematic exploration of these reasoning failures and identify a primary inducement: the latent spurious feature correlations within the input that drive these erroneous reasoning shortcuts. Driven by these findings, we introduce Causal-Contrastive Preference Optimization (C2PO), a unified alignment framework designed to tackle these specific failures by simultaneously discovering and suppressing these correlations directly within the optimization process. Specifically, C2PO leverages causal counterfactual signals to isolate bias-inducing features from valid reasoning paths, and employs a fairness-sensitive preference update mechanism to dynamically evaluate logit-level contributions and suppress shortcut features. Extensive experiments across multiple benchmarks covering stereotypical bias (BBQ, Unqover), structural bias (MNLI, HANS, Chatbot, MT-Bench), out-of-domain fairness (StereoSet, WinoBias), and general utility (MMLU, GSM8K) demonstrate that C2PO effectively mitigates stereotypical and structural biases while preserving robust general reasoning capabilities.

大模型偏见因果推理公平性优化偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。