arXiv:2608.05880cs.LGcs.AI2026-08中稿 · 36th Irish Signals…

对比三种方法在聚类模式检测中的表现,发现均无法全面识别所有模式。

Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation

论文配图:Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation
图 1 · 摘自论文原文
  • 用合成数据注入预设模式,系统评估解释方法
  • 三种方法都能找回相关特征,但无法全覆盖模式类型
  • 提示需专门设计模式检测工具,适合医疗等高维数据分析

聚类结果的可解释性仍是数据分析中的核心挑战,尤其在医疗等高维数据领域,需从复杂数据中提取有意义的结构化模式。现有可解释性技术多聚焦于特征重要性或局部实例解释,难以识别聚类内部的模式。本文对三种常用后处理分析方法进行对比评估:基于随机森林的置换特征重要性、LIME(局部可解释模型无关解释)和主成分分析。为实现可控评估,引入一组合成数据集,其中系统性注入预设模式。结果表明,尽管各方法均可成功恢复相关特征,但无一能一致检测所有注入的模式类型。该发现揭示了现有解释工具与模式级聚类解释需求之间的关键差距,推动发展专用的模式检测方法。

原文摘要 · Abstract (English)

Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they are primarily designed to assess feature importance or provide local instance-level explanations rather than to identify structured patterns present within clusters. This work presents a comparative evaluation of commonly used post-hoc analysis methods for pattern detection in clustering results. To enable controlled evaluation, we introduce a suite of synthetic datasets in which predefined patterns are systematically injected. Three widely used techniques are evaluated: a Random Forest surrogate model with permutation feature importance, LIME (Local Interpretable Model-agnostic Explanations), and principal component analysis. Results demonstrate that although each method can successfully recover relevant features, none consistently detects all injected pattern types. These findings high- light a critical gap between existing explainability tools and the requirements of pattern-level cluster interpretation, motivating the development of dedicated pattern detection methodologies.

聚类解释模式检测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。