少数群体通过集体重标注数据,低成本提升模型公平性。
Fairness for the People, by the People: Minority Collective Action
- 少数群体协作重标注自身数据,不改变算法流程。
- 小规模子群体行动可显著降低不公平性,误差增长有限。
- 适合关注公平性的用户或数据贡献者参考。
机器学习模型常继承训练数据中的偏见,导致对某些少数群体的不公平对待。尽管已有多种企业侧的偏见缓解技术,但通常伴随性能损失且需组织配合。鉴于许多模型依赖用户贡献的数据,我们提出通过算法集体行动框架,让少数群体协作战略性地重标注自身数据,从而提升公平性,且无需修改企业的训练过程。本文提出三种实用、模型无关的方法来近似理想重标注,并在真实数据集上验证。结果表明,少数群体中的小部分子群体即可显著减少不公平性,对整体预测误差影响很小。
原文摘要 · Abstract (English)
Machine learning models often preserve biases present in training data, leading to unfair treatment of certain minority groups. Despite an array of existing firm-side bias mitigation techniques, they typically incur utility costs and require organizational buy-in. Recognizing that many models rely on user-contributed data, end-users can induce fairness through the framework of Algorithmic Collective Action, where a coordinated minority group strategically relabels its own data to enhance fairness, without altering the firm's training process. We propose three practical, model-agnostic methods to approximate ideal relabeling and validate them on real-world datasets. Our findings show that a subgroup of the minority can substantially reduce unfairness with a small impact on the overall prediction error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。