EmbedOR通过曲率增强距离度量,更好保留高维数据聚类结构。
EmbedOR: Provable Cluster-Preserving Visualizations with Curvature-Based Stochastic Neighbor Embeddings
- 基于图曲率设计新距离度量,增强聚类结构保留能力
- 实验显示其极少分裂连续高密度区域,优于UMAP和tSNE
- 可用来标注可视化结果,识别碎片化并揭示数据几何
Stochastic Neighbor Embedding (SNE) 算法如 UMAP 和 tSNE 在处理噪声高维数据时,常无法保持数据的几何结构。具体表现为错误分离潜在数据子流形的连通分量,或在可良好聚类的数据中未能发现簇。为解决这些问题,我们提出 EmbedOR,一种融合离散图曲率的 SNE 算法。该算法使用曲率增强的距离度量进行随机嵌入,强调底层聚类结构。关键的是,我们证明该距离度量将 tSNE 的一致性结果推广至更广泛的数据库。我们在合成与真实数据上进行了大量实验,验证了 EmbedOR 在可视化与几何保持方面的优势。结果显示,与其他 SNE 算法及 UMAP 相比,EmbedOR 更少分裂连续、高密度的数据区域。此外,我们还证明该距离度量可用于标注现有可视化,识别碎片化现象,并提供对数据底层几何的更深层洞察。
原文摘要 · Abstract (English)
Stochastic Neighbor Embedding (SNE) algorithms like UMAP and tSNE often produce visualizations that do not preserve the geometry of noisy and high dimensional data. In particular, they can spuriously separate connected components of the underlying data submanifold and can fail to find clusters in well-clusterable data. To address these limitations, we propose EmbedOR, a SNE algorithm that incorporates discrete graph curvature. Our algorithm stochastically embeds the data using a curvature-enhanced distance metric that emphasizes underlying cluster structure. Critically, we prove that the EmbedOR distance metric extends consistency results for tSNE to a much broader class of datasets. We also describe extensive experiments on synthetic and real data that demonstrate the visualization and geometry-preservation capabilities of EmbedOR. We find that, unlike other SNE algorithms and UMAP, EmbedOR is much less likely to fragment continuous, high-density regions of the data. Finally, we demonstrate that the EmbedOR distance metric can be used as a tool to annotate existing visualizations to identify fragmentation and provide deeper insight into the underlying geometry of the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。