用拓扑分析定位GPT-2中导致偏见的注意力头
Augmenting Bias Detection in LLMs Using Topological Data Analysis
- 通过拓扑数据分析识别引发偏见的注意力头
- 发现性别、职业等偏见集中在特定注意力头热点区域
- 可精准定位特定群体的偏见来源,适合模型可解释性研究
近期许多偏见检测方法被提出以评估大语言模型的偏见程度,但尚未充分探索哪些模型组件导致对特定群体的偏见。本研究采用拓扑数据分析方法,识别GPT-2中影响StereoSet数据集中身份群体误代表的注意力头。结果发现,性别或职业等特定类别偏见集中在某些作为热点的注意力头中。所提出的度量方法还可用于确定某一偏见类别内特定群体的偏见来源,未来工作可将此方法拓展用于大模型去偏。
原文摘要 · Abstract (English)
Recently, many bias detection methods have been proposed to determine the level of bias a large language model captures. However, tests to identify which parts of a large language model are responsible for bias towards specific groups remain underdeveloped. In this study, we present a method using topological data analysis to identify which heads in GPT-2 contribute to the misrepresentation of identity groups present in the StereoSet dataset. We find that biases for particular categories, such as gender or profession, are concentrated in attention heads that act as hot spots. The metric we propose can also be used to determine which heads capture bias for a specific group within a bias category, and future work could extend this method to help de-bias large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。