arXiv:2509.08146cs.CLcs.LG2025-09EMNLP被引 4

提示词会传递大模型固有偏见,现有缓解方法效果不稳定。

Bias after Prompting: Persistent Discrimination in Large Language Models

  • 通过因果模型研究提示词适配中的偏见传播机制。
  • 性别、年龄、宗教偏见在提示后仍保持强相关(如性别rho≥0.94)。
  • 适用于关注模型公平性与提示工程的AI研发人员。

先前关于偏见迁移假说(BTH)的研究存在一个危险假设:预训练大语言模型(LLMs)的偏见不会传递到适配后的模型中。本文通过在因果模型下研究提示适配中的BTH,否定了这一假设。提示作为现实应用中广泛使用的适配策略,会导致偏见传播,且主流提示式缓解方法无法稳定阻止偏见转移。例如,在共指消解任务中,性别偏见相关性高达rho≥0.94;在问答任务中,年龄偏见rho≥0.98,宗教偏见rho≥0.69。即使改变少量样本组成参数(如样本量、刻板内容、职业分布、代表性平衡),偏见相关性仍维持在rho≥0.90以上。评估多种提示式去偏策略发现,不同方法各有优势,但无一能跨模型、任务和人群持续降低偏见传播。结果表明,仅修正模型内在偏见,可能不足以阻断其向下游任务的传播。

原文摘要 · Abstract (English)

A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapted models. We invalidate this assumption by studying the BTH in causal models under prompt adaptations, as prompting is an extremely popular and accessible adaptation strategy used in real-world applications. In contrast to prior work, we find that biases can transfer through prompting and that popular prompt-based mitigation methods do not consistently prevent biases from transferring. Specifically, the correlation between intrinsic biases and those after prompt adaptation remain moderate to strong across demographics and tasks -- for example, gender (rho >= 0.94) in co-reference resolution, and age (rho >= 0.98) and religion (rho >= 0.69) in question answering. Further, we find that biases remain strongly correlated when varying few-shot composition parameters, such as sample size, stereotypical content, occupational distribution and representational balance (rho >= 0.90). We evaluate several prompt-based debiasing strategies and find that different approaches have distinct strengths, but none consistently reduce bias transfer across models, tasks or demographics. These results demonstrate that correcting bias, and potentially improving reasoning ability, in intrinsic models may prevent propagation of biases to downstream tasks.

大模型偏见传播提示工程公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。