arXiv:2508.08855cs.CLcs.AI2025-08中稿 · EMNLP

通过注入偏差分析并消除大模型中的刻板印象。

BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection

  • 用标记微调向模型注入特定偏见,保持模型冻结。
  • 可定位偏见来源并精准抑制,不降低下游任务性能。
  • 适用于未见过的偏见类型,适合安全与可解释性研究。

理解大型语言模型(LLMs)权重中编码的偏见和刻板印象,对开发有效的缓解策略至关重要。然而,偏见行为往往微妙且难以分离,即使刻意诱发也难系统分析与去偏。为此,我们提出一个简单、低成本且通用的框架 exttt{BiasGym},用于可靠地注入、分析和缓解 LLM 内部的概念关联偏见。该框架包含两个模块: exttt{Inject} 通过基于标记的微调将特定偏见注入模型,同时保持模型冻结;随后是两种去偏方法,利用这些注入信号识别并可靠抑制( exttt{Scope})或引导( exttt{Steer})导致偏见行为的组件。该框架实现一致的偏见激发,有助于更精确地定位模型空间中的偏见概念关联,支持针对性去偏而无需牺牲下游任务表现,并可泛化至微调阶段未见的偏见。我们在减少真实世界刻板印象(如意大利人是‘鲁莽司机’)方面验证了该框架的有效性,展示了其在安全干预和可解释性研究中的实用价值。

原文摘要 · Abstract (English)

Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behavior is often subtle and non-trivial to isolate, even when deliberately elicited, making systematic analysis and debiasing particularly challenging. To address this, we introduce a simple, cost-effective, and generalizable framework \texttt{BiasGym} for reliably injecting, analyzing, and mitigating conceptual associations of biases within LLMs. \texttt{BiasGym} consists of two modules: \texttt{Inject}, which injects specific biases into the model via token-based fine-tuning while keeping the model frozen, followed by two debiasing methods that leverage these injected signals to identify and reliably suppress (\texttt{Scope}) or \texttt{Steer} the components responsible for biased behavior. Our framework enables consistent bias elicitation for better localization of bias conceptual association in the model space, supports targeted debiasing without degrading performance on downstream tasks, and generalizes to biases unseen during fine-tuning. We demonstrate the effectiveness of our proposed framework in reducing real-world stereotypes (e.g., people from Italy being `reckless drivers'), showing its utility for both safety interventions and interpretability research.

大模型去偏可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。