arXiv:2607.18119stat.MLcs.LG2026-07

提出新方法提升多敏感属性下的聚类公平性,计算更稳定高效。

COVAriance-Induced Fairness Gap Penalty for Subgroup-Fair Clustering

  • 用协方差构造公平性差距的连续近似,支持梯度优化
  • 在多个基准数据集上实现成本与公平性的良好平衡
  • 适用于多敏感属性子群体,适合关注算法公平性的研究者

公平聚类旨在使聚类分配独立于敏感属性,但当多个敏感属性共同定义大量子群体时,该目标面临挑战。此时直接扩展现有公平聚类算法会带来计算开销大或数值不稳定的难题,尤其当子群体数量呈指数增长且部分子群体样本极少时。为此,本文定义了聚类中的子群体公平性差距,并推导出一个与之精确匹配的协方差替代指标。进一步引入该指标的连续松弛形式,实现基于梯度的高效优化,提出 COVA-FC 算法。同时发现子群体公平性并不蕴含边际公平性,因此拓展框架以捕捉子群体-边际公平性差距。在多个基准数据集上的实验表明,COVA-FC 在子群体和高阶边际设置下均实现了具有竞争力的成本-公平性权衡,并显著优于现有基线的计算效率。

原文摘要 · Abstract (English)

Fair clustering aims to make cluster assignments independent of sensitive attributes, but this goal becomes challenging when multiple sensitive attributes jointly define many subgroups. In such settings, directly extending existing fair clustering algorithms is computationally expensive or numerically unstable, especially when the number of subgroups grows exponentially and some subgroups contain only a few instances. To address these challenges, we define a subgroup-fairness gap for clustering and derive a covariance-based surrogate that exactly matches this gap. We then introduce a continuous relaxation of the surrogate, enabling efficient gradient-based optimization and yielding our proposed algorithm, COVA-FC. We also show that subgroup fairness alone does not imply marginal fairness, and extend our framework to capture a subgroup-marginal-fairness gap. Experiments on benchmark datasets show that COVA-FC achieves competitive cost-fairness trade-offs and improves computational efficiency over existing baselines in both subgroup and higher-order marginal settings.

公平聚类子群体公平协方差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。