arXiv:2512.07249cs.LGcs.AI2025-12

用影响函数重加权样本,让分类模型更公平

IFFair: Influence Function-driven Sample Reweighting for Fair Classification

  • 基于样本对不同群体的影响差异动态调整权重
  • 在多个真实数据集上同时降低多种偏差指标
  • 无需改动模型结构,适合快速部署到现有系统

随着机器学习在社会决策中广泛应用,其基于数据的模式可能导致算法学习并加剧样本中的潜在偏见,从而对特定弱势群体造成歧视性决策,损害社会福祉并阻碍应用发展。为此,我们提出一种基于影响函数的预处理方法IFFair。与现有公平优化方法相比,IFFair仅利用训练样本对不同群体的影响差异作为指导,在不修改网络结构、数据特征和决策边界的情况下,动态调整训练过程中的样本权重。我们在多个真实数据集和评估指标上验证了该方法的有效性。实验结果表明,该方法在分类任务中显著降低了多种公平性指标的偏差,包括人口均等性、平等机会、错误率均等性等,且无冲突。同时,相比以往预处理方法,它在性能与公平性之间实现了更好的权衡。

原文摘要 · Abstract (English)

Because machine learning has significantly improved efficiency and convenience in the society, it's increasingly used to assist or replace human decision-making. However, the data-based pattern makes related algorithms learn and even exacerbate potential bias in samples, resulting in discriminatory decisions against certain unprivileged groups, depriving them of the rights to equal treatment, thus damaging the social well-being and hindering the development of related applications. Therefore, we propose a pre-processing method IFFair based on the influence function. Compared with other fairness optimization approaches, IFFair only uses the influence disparity of training samples on different groups as a guidance to dynamically adjust the sample weights during training without modifying the network structure, data features and decision boundaries. To evaluate the validity of IFFair, we conduct experiments on multiple real-world datasets and metrics. The experimental results show that our approach mitigates bias of multiple accepted metrics in the classification setting, including demographic parity, equalized odds, equality of opportunity and error rate parity without conflicts. It also demonstrates that IFFair achieves better trade-off between multiple utility and fairness metrics compared with previous pre-processing methods.

公平分类影响函数样本重加权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。