将公平性视为对称操作,通过正则化减少机器学习中的偏见。
Detecting and Mitigating Bias by Treating Fairness as a Symmetry Operation

- 把公平性建模为敏感属性翻转下的输出不变性
- 在四个数据集上偏见减少超90%,准确率损失仅5%左右
- 无需因果图知识,适用于多种敏感属性场景
部署于高风险社会经济场景的机器学习系统常表现出偏见。本文将偏见形式化为对称性破坏:若分类器在保持能力特征不变的情况下,对敏感属性进行反事实翻转后输出不变,则认为其公平。通过基于损失的正则化实现对称性恢复,在四个含不同噪声、相关性和偏见水平的合成数据集上评估,该框架实现超过90%的偏见违反降低,准确率代价约为5%。该方法无需因果图知识,计算轻量,且可推广至任意可表示为比特翻转的敏感属性,适用于主流基准中未包含局部歧视来源的场景。
原文摘要 · Abstract (English)
Machine learning systems deployed in high stakes socioeconomic settings routinely display bias. We formalize bias as a symmetry breaking operation: a classifier is fair if its outputs remain invariant under the counterfactual operation of switching a sensitive attribute, with merit features held fixed. We implement loss based regularization as a symmetry restoring mechanism and evaluate the framework on four synthetic datasets with varying levels of noise, correlation, and bias. The framework achieves upwards of 90\% violation reduction, with accuracy costs around 5\%. This framework does not require causal graph knowledge, is computationally lightweight, and generalizes to any sensitive attribute definable as a bit-flip, making it suitable for contexts where local sources of discrimination remain absent from mainstream benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。