用信息论方法自动分组高维数据嵌入,并生成简洁解释。
InfoClus: Informative Clustering of High-dimensional Data Embeddings
- 基于信息论构建可解释的聚类目标函数
- 在三个数据集上优于现有方法,尤其擅长发现生物特征模式
- 适合需要探索降维散点图的科研人员
理解高维数据可通过降维可视化实现,但低维嵌入难以解释。为此,我们提出带解释的划分概念:将嵌入中的数据划分为若干组,每组用原始高维属性生成稀疏解释。我们设计了一个基于信息论的目标函数,衡量从解释中能获取的知识量与解释复杂度之间的权衡。通过调节复杂度参数,可控制划分粒度。我们提出InfoClus,通过受层级聚类约束的贪心搜索联合优化划分与解释。在三个数据集上的定性与定量分析表明,与公开的手动分析结果(Cytometry数据)及两种近期嵌入解释方法(RVX和VERA)相比,InfoClus具有显著优势。结果表明,InfoClus能自动为降维散点图分析提供良好起点。
原文摘要 · Abstract (English)
Developing an understanding of high-dimensional data can be facilitated by visualizing that data using dimensionality reduction. However, the low-dimensional embeddings are often difficult to interpret. To facilitate the exploration and interpretation of low-dimensional embeddings, we introduce a new concept named partitioning with explanations. The idea is to partition the data shown through the embedding into groups, each of which is given a sparse explanation using the original high-dimensional attributes. We introduce an objective function that quantifies how much we can learn through observing the explanations of the data partitioning, using information theory, and also how complex the explanations are. Through parameterization of the complexity, we can tune the solutions towards the desired granularity. We propose InfoClus, which optimizes the partitioning and explanations jointly, through greedy search constrained over a hierarchical clustering. We conduct a qualitative and quantitative analysis of InfoClus on three data sets. We contrast the results on the Cytometry data with published manual analysis results, and compare with two other recent methods for explaining embeddings (RVX and VERA). These comparisons highlight that InfoClus has distinct advantages over existing procedures and methods. We find that InfoClus can automatically create good starting points for the analysis of dimensionality-reduction-based scatter plots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。