arXiv:2606.14592stat.MLcs.LG2026-06

提出可解释聚类结果的通用方法,判断哪些特征关键。

Cluster LOCO: Feature Importance For Interpreting Clusters

论文配图:Cluster LOCO: Feature Importance For Interpreting Clusters
图 1 · 摘自论文原文
  • 通过移除特征看聚类泛化能力下降程度,评估特征重要性。
  • 在合成数据和单细胞转录组中,比现有方法更准识别关键特征。
  • 不依赖特定算法,适合大规模数据,科研人员可信赖聚类结果。

聚类广泛应用于探索性分析与科学发现,从市场细分到生物数据分析,但随着数据规模和复杂度增加,其结果难以解释、审计与复现。可靠使用聚类需理解驱动结构的关键特征,然而相比监督学习,聚类的特征级解释仍很稀缺。现有方法常依赖特定算法与数据假设。为此,我们提出 Cluster LOCO(Leave-One-Covariate-Out),一种模型无关的聚类特征重要性评分方法。该方法基于特征遮蔽与聚类泛化能力,即在子集上学习的聚类标签能否准确预测保留样本。对任意聚类算法,通过测量移除某特征后泛化能力的下降程度来量化其重要性。我们先引入 Cluster LOCO-Split(基于数据分割),再扩展为适用于大规模数据的 Cluster LOCO-MP(小批次集成版本)。在合成模拟及单细胞转录组细胞类型发现应用中,结果表明 Cluster LOCO 比现有方法更可靠地恢复出有信息量的特征。

原文摘要 · Abstract (English)

Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex. Reliable use of clustering requires understanding which features drive the discovered structure, yet feature-level explanations for clustering remain scarce compared with methods in supervised learning. Furthermore, existing clustering feature importance scores are often tied to specific algorithms and data assumptions. To address these challenges, we propose Cluster LOCO (Leave-One-Covariate-Out), a family of model-agnostic feature importance scores for clustering. Cluster LOCO is built on feature occlusion and clustering generalizability, defined as whether cluster labels learned on one subset of the data can be accurately predicted on held-out samples. For any chosen clustering algorithm, Cluster LOCO quantifies a feature's importance by measuring how much its removal degrades generalizability. We first introduce Cluster LOCO-Split, which relies on data splitting, and then extend it to Cluster LOCO-MP, a minipatch ensemble-based version designed for large-scale data. Across synthetic simulations and an application to cell-type discovery in single-cell transcriptomics, we show that Cluster LOCO more reliably recovers informative features than existing clustering feature importance methods.

聚类解释特征重要性单细胞可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。