对比三种方法在聚类模式检测中的表现,发现均无法全面识别所有模式。
Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation

- 用合成数据注入预设模式,系统评估解释方法
- 三种方法都能找回相关特征,但无法全覆盖模式类型
- 提示需专门设计模式检测工具,适合医疗等高维数据分析
聚类结果的可解释性仍是数据分析中的核心挑战,尤其在医疗等高维数据领域,需从复杂数据中提取有意义的结构化模式。现有可解释性技术多聚焦于特征重要性或局部实例解释,难以识别聚类内部的模式。本文对三种常用后处理分析方法进行对比评估:基于随机森林的置换特征重要性、LIME(局部可解释模型无关解释)和主成分分析。为实现可控评估,引入一组合成数据集,其中系统性注入预设模式。结果表明,尽管各方法均可成功恢复相关特征,但无一能一致检测所有注入的模式类型。该发现揭示了现有解释工具与模式级聚类解释需求之间的关键差距,推动发展专用的模式检测方法。
原文摘要 · Abstract (English)
Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they are primarily designed to assess feature importance or provide local instance-level explanations rather than to identify structured patterns present within clusters. This work presents a comparative evaluation of commonly used post-hoc analysis methods for pattern detection in clustering results. To enable controlled evaluation, we introduce a suite of synthetic datasets in which predefined patterns are systematically injected. Three widely used techniques are evaluated: a Random Forest surrogate model with permutation feature importance, LIME (Local Interpretable Model-agnostic Explanations), and principal component analysis. Results demonstrate that although each method can successfully recover relevant features, none consistently detects all injected pattern types. These findings high- light a critical gap between existing explainability tools and the requirements of pattern-level cluster interpretation, motivating the development of dedicated pattern detection methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。