用SHAP值加权特征,让无监督聚类效果提升22.69%。
Refining Filter Global Feature Weighting for Fully-Unsupervised Clustering
- 用SHAP值计算特征重要性,动态加权用于聚类
- 在5个数据集上使调整兰德指数最高提升22.69%
- 适合追求高精度聚类且需解释性的研究者
在无监督学习中,有效聚类对揭示未标记数据中的模式至关重要。然而,聚类性能常依赖于特征的相关性和贡献度,而这些差异在不同数据集间显著。本文探索聚类中的特征加权策略,提出基于SHAP(SHapley Additive exPlanations)的新方法。不同于传统仅用于可解释性,本文将SHAP值用于特征加权,直接提升无监督聚类质量。在五个基准数据集和多种聚类方法上的实证评估表明,基于SHAP的特征加权能显著改善聚类效果,调整兰德指数从0.586提升至0.719,最高提升达22.69%。文中还深入分析了加权带来增益的具体场景,为实际应用提供指导。
原文摘要 · Abstract (English)
In the context of unsupervised learning, effective clustering plays a vital role in revealing patterns and insights from unlabeled data. However, the success of clustering algorithms often depends on the relevance and contribution of features, which can differ between various datasets. This paper explores feature weighting for clustering and presents new weighting strategies, including methods based on SHAP (SHapley Additive exPlanations), a technique commonly used for providing explainability in various supervised machine learning tasks. By taking advantage of SHAP values in a way other than just to gain explainability, we use them to weight features and ultimately improve the clustering process itself in unsupervised scenarios. Our empirical evaluations across five benchmark datasets and clustering methods demonstrate that feature weighting based on SHAP can enhance unsupervised clustering quality, achieving up to a 22.69\% improvement over other weighting methods (from 0.586 to 0.719 in terms of the Adjusted Rand Index). Additionally, these situations where the weighted data boosts the results are highlighted and thoroughly explored, offering insight for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。