政府数据中的偏见难以通过常规方法消除,根源在数据本身的历史与结构。
Failing on Bias Mitigation: A Case Study on the Challenges of Fairness in Government Data
- 用布里斯托尔市犯罪预测数据测试多种公平性缓解方法
- 所有方法均无法克服数据中固有的不公平性,因数据本身含历史偏见
- 提醒政策制定者:仅靠算法优化无法解决系统性偏见问题
AI支持的政府服务存在偏见和不公风险,引发伦理与法律担忧。以布里斯托尔市议会提供的犯罪率预测数据为例,我们研究了为何广泛采用的偏见缓解技术在政府数据上常失效。不同于对实际部署系统的审计,本研究旨在揭示这些方法失败的根本原因。实验对比多种模型与公平性技术,发现即便在模型架构与评估指标合理的情况下,缓解措施仍无法克服数据中嵌入的不公平性——其根源在于政府数据自身的结构与历史。我们进一步分析了数据分布变化、历史偏见累积及数据延迟发布等导致失败的原因。此外,在多敏感特征交叉实验中,发现现有公平性分析存在盲区。尽管研究局限于单一城市,但结果具有警示意义:标准缓解方法可能无法有效应对政府数据中的深层偏见。
原文摘要 · Abstract (English)
The potential for bias and unfairness in AI-supporting government services raises ethical and legal concerns. Using crime rate prediction with the Bristol City Council data as a case study, we examine how these issues persist. Rather than auditing real-world deployed systems, our goal is to understand why widely adopted bias mitigation techniques often fail when applied to government data. Our findings reveal that bias mitigation approaches applied to government data are not always effective -- not because of flaws in model architecture or metric selection, but due to the inherent properties of the data itself. Through comparing a set of comprehensive models and fairness methods, our experiments consistently show that the mitigation efforts cannot overcome the embedded unfairness in the data -- further reinforcing that the origin of bias lies in the structure and history of government datasets. We then explore the reasons for the mitigation failures in predictive models on government data and highlight the potential sources of unfairness posed by data distribution shifts, the accumulation of historical bias, and delays in data release. We also discover the limitations of the blind spots in fairness analysis and bias mitigation methods when only targeting a single sensitive feature through a set of intersectional fairness experiments. Although this study is limited to one city, the findings are highly suggestive, which can contribute to an early warning that biases in government data may persist even with standard mitigation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。