arXiv:2504.01030stat.MLcs.LG2025-04

平衡公平与信息保留,提升模型决策公正性

Fair Sufficient Representation Learning

  • 用凸组合统一优化充分性与公平性目标
  • 在医疗和文本数据上实现更高公平-准确权衡
  • 适合需避免偏见的高风险决策场景

公平统计建模与机器学习的核心目标是降低或消除数据或模型自身带来的偏差,确保预测与决策不受种族、性别、年龄等敏感属性的不公正影响。本文提出公平充分表示学习(FSRL)方法,平衡表示的充分性与公平性。充分性要求表示包含目标变量的所有必要信息,公平性则要求表示与敏感属性独立。FSRL基于充分性目标与公平性目标的凸组合,通过距离协方差刻画随机变量间的独立性,在表示层面实现公平与充分的协同优化。我们建立了所学表示的收敛性理论,并在具有不同结构的健康和文本数据集上进行实验,结果表明,相比现有方法,FSRL在公平性与准确率之间实现了更优的权衡。

原文摘要 · Abstract (English)

The main objective of fair statistical modeling and machine learning is to minimize or eliminate biases that may arise from the data or the model itself, ensuring that predictions and decisions are not unjustly influenced by sensitive attributes such as race, gender, age, or other protected characteristics. In this paper, we introduce a Fair Sufficient Representation Learning (FSRL) method that balances sufficiency and fairness. Sufficiency ensures that the representation should capture all necessary information about the target variables, while fairness requires that the learned representation remains independent of sensitive attributes. FSRL is based on a convex combination of an objective function for learning a sufficient representation and an objective function that ensures fairness. Our approach manages fairness and sufficiency at the representation level, offering a novel perspective on fair representation learning. We implement this method using distance covariance, which is effective for characterizing independence between random variables. We establish the convergence properties of the learned representations. Experiments conducted on healthcase and text datasets with diverse structures demonstrate that FSRL achieves a superior trade-off between fairness and accuracy compared to existing approaches.

公平学习表示学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。