提出一种可自动确定聚类数的公平贝叶斯聚类方法
Fair Bayesian Model-Based Clustering

- 设计专用先验使聚类结果天然满足群体公平性
- 能自动推断聚类数量且适用于各类数据类型
- 在真实数据上兼顾公平性与聚类质量,尤其适合分类数据
随着机器学习技术的发展和对可信AI的需求增加,公平聚类已成为重要的社会议题。群体公平要求各敏感群体在所有聚类中的比例相似。现有大多数公平聚类方法基于K-means,需预先设定距离度量和聚类数量。为解决此限制,本文提出公平贝叶斯模型聚类(Fair Bayesian Clustering, FBC),设计仅在公平聚类上有质量的先验,并实现高效的MCMC算法。FBC的优势在于可自动推断聚类数量,且只要定义了似然函数即可应用于任意数据类型(如分类数据)。在真实数据集上的实验表明,FBC (i) 能合理推断聚类数量,(ii) 在效用-公平性权衡上表现优于现有方法,(iii) 在分类数据上表现良好。
原文摘要 · Abstract (English)
Fair clustering has become a socially significant task with the advancement of machine learning technologies and the growing demand for trustworthy AI. Group fairness ensures that the proportions of each sensitive group are similar in all clusters. Most existing group-fair clustering methods are based on the $K$-means clustering and thus require the distance between instances and the number of clusters to be given in advance. To resolve this limitation, we propose a fair Bayesian model-based clustering called Fair Bayesian Clustering (FBC). We develop a specially designed prior which puts its mass only on fair clusters, and implement an efficient MCMC algorithm. Advantages of FBC are that it can infer the number of clusters and can be applied to any data type as long as the likelihood is defined (e.g., categorical data). Experiments on real-world datasets show that FBC (i) reasonably infers the number of clusters, (ii) achieves a competitive utility-fairness trade-off compared to existing fair clustering methods, and (iii) performs well on categorical data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。