首个可解释的聚类集成方法,让算法决策透明可信。
Interpretable Clustering Ensemble
- 将基聚类结果当类别变量,原空间建决策树
- 性能媲美顶尖方法,同时具备可解释性
- 适合医疗、金融等需透明决策的场景
聚类集成已成为机器学习的重要研究方向。尽管已有众多方法提升聚类质量,但多数现有方法忽视了高风险应用中的可解释性需求。在医疗诊断和金融风控等领域,算法不仅需准确,还需可解释以保障决策透明可信。为此,我们提出文献中首个专为聚类集成设计的可解释算法。通过将基聚类划分视为分类变量,在原始特征空间构建决策树,并利用统计关联检验指导树的生成过程。实验表明,该算法在性能上与当前最优(SOTA)聚类集成方法相当,同时具备额外的可解释性。据我们所知,这是首个专门针对聚类集成的可解释算法,为可解释聚类研究提供了新视角。
原文摘要 · Abstract (English)
Clustering ensemble has emerged as an important research topic in the field of machine learning. Although numerous methods have been proposed to improve clustering quality, most existing approaches overlook the need for interpretability in high-stakes applications. In domains such as medical diagnosis and financial risk assessment, algorithms must not only be accurate but also interpretable to ensure transparent and trustworthy decision-making. Therefore, to fill the gap of lack of interpretable algorithms in the field of clustering ensemble, we propose the first interpretable clustering ensemble algorithm in the literature. By treating base partitions as categorical variables, our method constructs a decision tree in the original feature space and use the statistical association test to guide the tree building process. Experimental results demonstrate that our algorithm achieves comparable performance to state-of-the-art (SOTA) clustering ensemble methods while maintaining an additional feature of interpretability. To the best of our knowledge, this is the first interpretable algorithm specifically designed for clustering ensemble, offering a new perspective for future research in interpretable clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。