arXiv:2606.20461cs.LGcs.CY2026-06中稿 · FAccT 2026被引 1

在数据有限时,用约束保证交叉群体代表性,平衡公平与效率。

Data Bias Mitigation under Coverage Constraints & The Price of Fairness

  • 引入覆盖约束,强制训练数据中包含交叉群体
  • 通过整数线性规划优化,实现低偏差高效率
  • 量化公平代价,帮助决策者权衡数据成本

机器学习模型在多个敏感属性(如种族与性别)交集群体上常出现歧视性结果或性能下降,原因在于缺乏可量化的偏见度量方法以及交集子群体在训练数据中的代表性不足。本文扩展了近期的偏见缓解框架,引入覆盖约束以确保各群体(包括交集子群体)有足够代表。由于对所有群体实现完全零偏见可能数据效率低下,本方案允许微小的偏见近似误差,换取更高的数据使用效率。我们还将偏见缓解建模为整数线性规划问题,刻画公平代价——即最小数据修改成本——作为公平容忍度的函数。该方法对法律合规(如监管要求特定公平阈值)和数据治理至关重要,使从业者能理性权衡偏见降低与数据修改(尤其是数据采购)成本之间的关系。我们在公开数据集上评估了该方法,结果显示:在多种分类器上,该框架在保持预测准确性的前提下有效缓解偏见;覆盖约束虽源于统计考量,却是维持下游模型性能的关键。

原文摘要 · Abstract (English)

Machine learning models have been shown to exhibit discriminatory outcomes or degraded performance for individuals at the intersection of multiple sensitive attributes, such as race and gender. This stems in part from two interrelated challenges: the lack of principled measures for quantifying bias (potentially intersectional), and insufficient representation of intersectional subgroups in training data. We extend a recent bias mitigation framework to incorporate coverage constraints that enforce sufficient representation across groups, including intersectional subgroups. Since achieving exactly zero bias for all groups may not be data efficient (meaning it may require large amounts of data), our solution trades small approximation errors in bias for greater data efficiency while satisfying coverage constraints. We also formulate bias mitigation as an integer linear program that optimizes over all mitigation strategies, and characterize the price of fairness, the minimum data modification cost, as a function of fairness tolerance. This is essential both for legal compliance, where regulations may mandate specific fairness thresholds, and for data governance, enabling practitioners to make informed trade-offs between bias reduction and data modification (particularly, data purchasing) costs. We evaluate our techniques on publicly available datasets, demonstrating that bias mitigation via our framework preserves predictive accuracy across multiple classifiers, and that coverage constraints, while motivated by statistical considerations, are essential for preserving downstream ML performance.

偏见缓解公平性数据效率覆盖约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。