arXiv:2606.04777cs.LG2026-06

统一优化公平性与聚类质量,减少群体间差异。

UniFair: A unified fair clustering approach based on separation and compactness

论文配图:UniFair: A unified fair clustering approach based on separation and compactness
图 1 · 摘自论文原文
  • 联合优化分离公平与社会公平,兼顾边界距离与组内代价
  • 在多个数据集上显著降低群体间差异,聚类损失仅小幅上升
  • 适用于需要公平聚类的医疗、金融等高风险决策场景

聚类正越来越多地用于支持高影响决策,但标准目标如k-means可能产生对不同人口群体不平等的聚类结果。现有公平聚类方法通常只优化单一公平概念,且常忽略聚类代价与决策边界几何结构之间的相互作用。本文提出UniFair,一种统一框架,同时优化分离公平性与社会公平性:分离公平性促使保护群体远离诱导的决策边界,社会公平性则通过惩罚组别级聚类代价来减少组内畸变差异。我们为分离公平和统一k-means目标开发了基于梯度的优化算法,并将其扩展至深度聚类,在自编码器的潜在空间中施加相同准则。在表格和图像数据集上的实验表明,UniFair在仅带来轻微聚类损失增加的前提下,有效降低了与边界相关的以及基于代价的群体差异。

原文摘要 · Abstract (English)

Clustering is increasingly used to support high-impact decisions, yet standard objectives such as k-means can produce clusterings that treat demographic groups unequally. Existing fair clustering methods typically optimize a single notion of fairness and often overlook how clustering costs interact with the geometry of the induced decision boundaries. We propose UniFair, a unified framework that jointly optimizes separation fairness and social fairness. Separation fairness encourages protected groups to lie farther from the induced decision boundaries, while social fairness reduces disparities in within-cluster distortion by penalizing group-wise clustering costs. We develop gradient-based optimization procedures for separation-fair and unified k-means objectives, and extend them to deep clustering by enforcing the same criteria in the latent space of an autoencoder. Experiments on tabular and image datasets show that UniFair reduces both boundary-related and cost-based group disparities with only a modest increase in clustering loss.

公平聚类k-means深度聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。