用简单聚类单元提升聚类公平性,兼顾效果与效率
Fair Clustering with Clusterlets
- 基于聚类单元匹配,优化公平性与经典聚类目标
- 参数调优后实现高凝聚性与低重叠度
- 适合关注算法公平性的实际应用开发者
随着聚类方法在现实中的广泛应用,其公平性成为重要研究方向。理论研究表明,公平性具有传递性:若存在若干小而公平的聚类,则通过简单的中心点聚类算法即可获得整体公平的聚类结果。然而,找到合适的初始聚类往往计算成本高、过程复杂或缺乏依据。本文提出一系列基于聚类单元(clusterlet)的模糊聚类算法,通过匹配单类聚类单元,优化公平聚类。匹配过程利用聚类单元距离,在优化经典聚类目标的同时引入公平性正则化。实验表明,简单的匹配策略即可实现高公平性,适当参数调优能同时达成高凝聚性与低重叠。
原文摘要 · Abstract (English)
Given their widespread usage in the real world, the fairness of clustering methods has become of major interest. Theoretical results on fair clustering show that fairness enjoys transitivity: given a set of small and fair clusters, a trivial centroid-based clustering algorithm yields a fair clustering. Unfortunately, discovering a suitable starting clustering can be computationally expensive, rather complex or arbitrary. In this paper, we propose a set of simple \emph{clusterlet}-based fuzzy clustering algorithms that match single-class clusters, optimizing fair clustering. Matching leverages clusterlet distance, optimizing for classic clustering objectives, while also regularizing for fairness. Empirical results show that simple matching strategies are able to achieve high fairness, and that appropriate parameter tuning allows to achieve high cohesion and low overlap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。