提出新方法让大模型在不同表述下保持公平一致性
DeFrame: Debiasing Large Language Models Against Framing Effects
- 引入'表述差异'量化框架对公平性的影响
- 发现现有去偏方法无法减少表述差异导致的偏差
- 新方法提升模型在不同表述下的公平性与一致性
随着大语言模型(LLMs)在真实场景中广泛应用,确保其在不同人群中的公平响应变得至关重要。尽管已有诸多努力,但隐性偏差问题仍存在:模型在标准评估中表现公平,但在非标准设置下可能产生偏差。本文识别出‘表述’(framing)——即语义等价提示的不同表达方式(如“A优于B” vs “B不如A”)——是造成这一差距的重要但未被充分研究的因素。我们首次提出‘表述差异’概念,用于量化表述对公平性评估的影响。通过在公平性基准上加入替代表述,发现:(1) 公平性评分随表述显著变化;(2) 现有去偏方法虽能提升整体(即平均帧)公平性,却常无法缓解由表述引发的差异。为此,我们提出一种面向表述的去偏方法,鼓励模型在不同表述间保持一致性。实验表明,该方法同时降低整体偏差并增强对表述差异的鲁棒性,使模型生成更公平、更一致的回应。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has become crucial. Despite many efforts, an ongoing challenge is hidden bias: LLMs appear fair under standard evaluations, but can produce biased responses outside those evaluation settings. In this paper, we identify framing -- differences in how semantically equivalent prompts are expressed (e.g., "A is better than B" vs. "B is worse than A") -- as an underexplored contributor to this gap. We first introduce the concept of "framing disparity" to quantify the impact of framing on fairness evaluation. By augmenting fairness evaluation benchmarks with alternative framings, we find that (1) fairness scores vary significantly with framing and (2) existing debiasing methods improve overall (i.e., frame-averaged) fairness, but often fail to reduce framing-induced disparities. To address this, we propose a framing-aware debiasing method that encourages LLMs to be more consistent across framings. Experiments demonstrate that our approach reduces overall bias and improves robustness against framing disparities, enabling LLMs to produce fairer and more consistent responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。