arXiv:2503.05371cs.LGcs.AI2025-03Conference of the …被引 17

用方向向量调控模型激活,有效降低大模型偏见。

Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs

  • 通过训练数据计算8个社会偏见方向向量,动态调节模型输出。
  • 在BBQ等4个数据集上平均提升12.8%,对主流方法有全面优势。
  • 计算高效且对通用能力影响最小,适合部署于实际系统。

我们提出一种新型大语言模型偏见缓解方法,通过在前向传播中应用方向向量调整模型激活。在BBQ数据集的子集上计算了8个对应不同社会偏见轴(如年龄、性别、种族)的方向向量,并在4个数据集上与3种其他方法对比。在BBQ数据集上优化后,各方向向量平均提升12.8%(BBQ)、8.3%(CLEAR-Bias)和1%(StereoSet),在所有测试中优于提示法和Self-Debias,且在17次评估中有12次优于微调。此外,方向向量对MMLU得分的影响最小。本工作首次系统性研究方向向量用于偏见缓解,证明其是强大且计算高效的策略,对提升AI安全性具有广泛意义。

原文摘要 · Abstract (English)

We present a novel approach to bias mitigation in large language models (LLMs) by applying steering vectors to modify model activations in forward passes. We compute 8 steering vectors, each corresponding to a different social bias axis, such as age, gender, or race, on a training subset of the BBQ dataset and compare the effectiveness of these to 3 additional bias mitigation methods across 4 datasets. When optimized on the BBQ dataset, our individually tuned steering vectors achieve average improvements of 12.8% on BBQ, 8.3% on CLEAR-Bias, and 1% on StereoSet, and show improvements over prompting and Self-Debias in all cases, and improvements over fine-tuning in 12 out of 17 evaluations. In addition, steering vectors showed the lowest impact on MMLU scores of the four bias mitigation methods tested. The work presents the first systematic investigation of steering vectors for bias mitigation, and we demonstrate that they are a powerful and computationally efficient strategy for reducing bias in LLMs, with broader implications for enhancing AI safety.

偏见缓解方向向量LLM安全高效方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。