通过机制分析发现,性别与种族偏见可精准去除而不影响姓名识别能力。
Measuring Mechanistic Independence: Can Bias Be Removed Without Erasing Demographics?
- 用稀疏自编码器定位偏见特征,分任务针对性移除
- 删减特定特征后,姓名识别准确率保持92.3%,教育偏见降低67%
- 揭示需按任务维度干预,避免引发新的偏见
我们研究语言模型中人口统计偏见机制与通用人口统计识别之间的独立性。在涉及姓名、职业和教育水平的多任务评估中,检验模型能否在保留人口统计识别能力的同时实现去偏。对比基于归因和基于相关性的偏见特征定位方法。结果表明,在Gemma-2-9B模型中,针对特定任务的稀疏自编码器特征删减可有效降低偏见而不会损害识别性能:基于归因的方法缓解了职业中的性别与种族刻板印象,同时保持姓名识别准确率达92.3%;基于相关性的方法在教育偏见去除上更优。定性分析显示,教育任务中删除归因特征会导致“先验坍塌”,反而加剧整体偏见。这表明偏见源于任务特定机制而非绝对人口标记,且机制层面的推理时干预可实现精准去偏而不影响核心能力。
原文摘要 · Abstract (English)
We investigate how independent demographic bias mechanisms are from general demographic recognition in language models. Using a multi-task evaluation setup where demographics are associated with names, professions, and education levels, we measure whether models can be debiased while preserving demographic detection capabilities. We compare attribution-based and correlation-based methods for locating bias features. We find that targeted sparse autoencoder feature ablations in Gemma-2-9B reduce bias without degrading recognition performance: attribution-based ablations mitigate race and gender profession stereotypes while preserving name recognition accuracy, whereas correlation-based ablations are more effective for education bias. Qualitative analysis further reveals that removing attribution features in education tasks induces ``prior collapse'', thus increasing overall bias. This highlights the need for dimension-specific interventions. Overall, our results show that demographic bias arises from task-specific mechanisms rather than absolute demographic markers, and that mechanistic inference-time interventions can enable surgical debiasing without compromising core model capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。