arXiv:2506.07861cs.LGcs.AI2025-06ICML被引 3

提出信息论框架,分析公平性过拟合问题并给出可落地的改进方向。

Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective

  • 用信息论中的互信息和条件互信息构建公平性泛化边界。
  • 理论推导出紧致的公平性泛化误差上界,验证其在多种算法中有效。
  • 为设计能泛化公平性的机器学习算法提供新思路,适合关注公平性研究者。

尽管在高风险应用中使用机器学习提升公平性已取得显著进展,但现有方法通常通过正则化等干预训练过程实现,缺乏对训练中获得的公平性能否推广到未见数据的严格保证。虽然预测性能的过拟合已被广泛研究,但公平性损失的过拟合却少有关注。本文从信息论视角提出一个分析公平性泛化误差的理论框架。基于Efron-Stein不等式,我们提出新颖的边界技术,推导出包含互信息(MI)和条件互信息(CMI)的紧致信息论公平性泛化上界。实验结果表明,这些边界在多种公平性感知学习算法中均具有紧致性和实际相关性。该框架为改进公平性泛化能力的算法设计提供了重要洞见。

原文摘要 · Abstract (English)

Despite substantial progress in promoting fairness in high-stake applications using machine learning models, existing methods often modify the training process, such as through regularizers or other interventions, but lack formal guarantees that fairness achieved during training will generalize to unseen data. Although overfitting with respect to prediction performance has been extensively studied, overfitting in terms of fairness loss has received far less attention. This paper proposes a theoretical framework for analyzing fairness generalization error through an information-theoretic lens. Our novel bounding technique is based on Efron-Stein inequality, which allows us to derive tight information-theoretic fairness generalization bounds with both Mutual Information (MI) and Conditional Mutual Information (CMI). Our empirical results validate the tightness and practical relevance of these bounds across diverse fairness-aware learning algorithms. Our framework offers valuable insights to guide the design of algorithms improving fairness generalization.

公平性信息论过拟合泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。