用梯度分析神经网络中的社会偏见,可定位并消除模型偏见。
GRADIEND: Feature Learning within Neural Networks Exemplified through Biases
- 通过模型梯度识别与性别、种族等相关的特征神经元。
- 能精准修改特定权重,实现去偏同时保留原有能力。
- 适用于多种模型结构,适合研究模型公平性的人参考。
人工智能系统常表现出并放大社会偏见,导致在关键领域产生有害后果。本研究提出一种新的编码器-解码器方法,利用模型梯度学习特征神经元中编码的社会偏见信息,如性别、种族和宗教。结果表明,该方法不仅能识别需调整的模型权重以改变特定特征,还能用于重写模型以实现去偏,同时保持其他功能不变。我们在多种模型架构上验证了该方法的有效性,展示了其广泛的应用潜力。
原文摘要 · Abstract (English)
AI systems frequently exhibit and amplify social biases, leading to harmful consequences in critical areas. This study introduces a novel encoder-decoder approach that leverages model gradients to learn a feature neuron encoding societal bias information such as gender, race, and religion. We show that our method can not only identify which weights of a model need to be changed to modify a feature, but even demonstrate that this can be used to rewrite models to debias them while maintaining other capabilities. We demonstrate the effectiveness of our approach across various model architectures and highlight its potential for broader applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。