arXiv:2411.05648cs.LG2024-11

通过相似性网络提升模型公平性与准确率,揭示数据偏差与分类复杂度的关系。

Enhancing Model Fairness and Accuracy with Similarity Networks: A Methodological Approach

  • 将数据映射到相似性特征空间,动态调节成对相似度分辨率
  • 实验验证方法在分类、数据补全和增强中均提升公平性表现
  • 适合关注模型公平性、数据偏见修复的研究者使用

本文提出一种新方法,深入探索下游机器学习任务中引入偏差的数据集特征。根据数据格式,采用不同技术将样本映射到相似性特征空间。该方法通过调整成对相似性的分辨率,清晰揭示了数据集分类复杂度与模型公平性之间的关系。实验结果证实,相似性网络在促进公平模型方面具有良好的应用前景。此外,该方法不仅在分类等下游任务中表现优异,还能有效实现满足人口均衡性等公平标准的数据集补全与增强。

原文摘要 · Abstract (English)

In this paper, we propose an innovative approach to thoroughly explore dataset features that introduce bias in downstream machine-learning tasks. Depending on the data format, we use different techniques to map instances into a similarity feature space. Our method's ability to adjust the resolution of pairwise similarity provides clear insights into the relationship between the dataset classification complexity and model fairness. Experimental results confirm the promising applicability of the similarity network in promoting fair models. Moreover, leveraging our methodology not only seems promising in providing a fair downstream task such as classification, it also performs well in imputation and augmentation of the dataset satisfying the fairness criteria such as demographic parity and imbalanced classes.

模型公平性相似性网络数据偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。