定位大模型中的刻板印象激活点,为消除偏见提供新思路。
Can We Locate and Prevent Stereotypes in LLMs?

- 通过对比神经元识别刻板印象相关激活
- 发现注意力头对偏差输出有显著贡献
- 首次在GPT-2和Llama中绘制偏见指纹图谱
大语言模型中的刻板印象可能加剧社会偏见。尽管模型应用广泛,但其内部偏见的存储位置仍不明确。本研究分析GPT-2 Small与Llama 3.2的内部机制,探索两种方法:识别编码刻板印象的对比神经元激活,以及检测对偏差输出贡献大的注意力头。实验旨在定位这些“偏见指纹”,为缓解刻板印象提供初步依据。
原文摘要 · Abstract (English)
Stereotypes in large language models (LLMs) can perpetuate harmful societal biases. Despite the widespread use of models, little is known about where these biases reside in the neural network. This study investigates the internal mechanisms of GPT 2 Small and Llama 3.2 to locate stereotype related activations. We explore two approaches: identifying individual contrastive neuron activations that encode stereotypes, and detecting attention heads that contribute heavily to biased outputs. Our experiments aim to map these "bias fingerprints" and provide initial insights for mitigating stereotypes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。