arXiv:2507.12021stat.MLcs.LG2025-07被引 1

让数据模型更公平:在原型分析中加入公平性约束。

Incorporating Fairness Constraints into Archetypal Analysis

  • 在原型分析中引入公平性正则项,降低敏感属性影响。
  • 在合成与真实数据上均显著降低群体可区分性,保持解释性。
  • 适合需要公平性保障的医疗、金融等敏感领域应用。

原型分析(AA)是一种无监督学习方法,将数据表示为极端模式(原型)的凸组合。尽管AA能提供可解释且低维的表示,但可能无意中编码敏感属性,引发公平性问题。本文提出公平原型分析(FairAA),通过显式减少敏感组信息对投影结果的影响来改进。我们还引入了非线性扩展FairKernelAA,以处理更复杂的数据分布。该方法在保留原型结构和可解释性的前提下,加入公平性正则化项。我们在合成数据集(包括线性、非线性及多组场景)上评估,结果显示其显著降低了群体可分性(以最大均值差异和线性可分性衡量),同时未明显牺牲解释方差。在真实世界ANSUR I数据集上的验证进一步证明了方法的鲁棒性和实用性。结果表明,FairAA在效用与公平性之间取得了良好平衡,是敏感应用场景中负责任表征学习的有力工具。

原文摘要 · Abstract (English)

Archetypal Analysis (AA) is an unsupervised learning method that represents data as convex combinations of extreme patterns called archetypes. While AA provides interpretable and low-dimensional representations, it can inadvertently encode sensitive attributes, leading to fairness concerns. In this work, we propose Fair Archetypal Analysis (FairAA), a modified formulation that explicitly reduces the influence of sensitive group information in the learned projections. We also introduce FairKernelAA, a nonlinear extension that addresses fairness in more complex data distributions. Our approach incorporates a fairness regularization term while preserving the structure and interpretability of the archetypes. We evaluate FairAA and FairKernelAA on synthetic datasets, including linear, nonlinear, and multi-group scenarios, demonstrating their ability to reduce group separability -- as measured by mean maximum discrepancy and linear separability -- without substantially compromising explained variance. We further validate our methods on the real-world ANSUR I dataset, confirming their robustness and practical utility. The results show that FairAA achieves a favorable trade-off between utility and fairness, making it a promising tool for responsible representation learning in sensitive applications.

原型分析公平性无监督学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。