统一聚类与解释,让混合类型数据聚类结果可懂可信赖。
Weight-Informed Self-Explaining Clustering for Mixed-Type Tabular Data
- 用稀疏编码对齐数值与类别特征,统一表示空间。
- 通过留一特征策略生成多视角加权,提升聚类质量。
- 输出可加性分解的特征级解释,适合需要透明决策的场景。
混合类型表格数据聚类是探索性分析的基础,但因数值与类别特征表示不一致、特征重要性分布不均且上下文依赖强、聚类与解释脱节而难以实现。本文提出 WISE 框架,将表示学习、特征加权、聚类和解释整合为全无监督、透明的统一流程。WISE 引入带填充的二值编码(BEP)将异构特征映射至统一稀疏空间;采用留一特征外(LOFO)策略获取多组高质量、多样化的特征加权视图;设计两阶段加权聚类过程,聚合多种语义划分。为保障内在可解释性,进一步提出判别频项(DFI),实现从实例到聚类的一致性特征级解释,并满足可加性分解保证。在六个真实数据集上的实验表明,WISE 在聚类质量上持续优于经典与神经基线方法,同时保持高效,并生成与聚类驱动原理一致的忠实、人类可读解释。
原文摘要 · Abstract (English)
Clustering mixed-type tabular data is fundamental for exploratory analysis, yet remains challenging due to misaligned numerical-categorical representations, uneven and context-dependent feature relevance, and disconnected and post-hoc explanation from the clustering process. We propose WISE, a Weight-Informed Self-Explaining framework that unifies representation, feature weighting, clustering, and interpretation in a fully unsupervised and transparent pipeline. WISE introduces Binary Encoding with Padding (BEP) to align heterogeneous features in a unified sparse space, a Leave-One-Feature-Out (LOFO) strategy to sense multiple high-quality and diverse feature-weighting views, and a two-stage weight-aware clustering procedure to aggregate alternative semantic partitions. To ensure intrinsic interpretability, we further develop Discriminative FreqItems (DFI), which yields feature-level explanations that are consistent from instances to clusters with an additive decomposition guarantee. Extensive experiments on six real-world datasets demonstrate that WISE consistently outperforms classical and neural baselines in clustering quality while remaining efficient, and produces faithful, human-interpretable explanations grounded in the same primitives that drive clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。