arXiv:2510.15014stat.MLcs.LG2025-10

提出t-SNE树,用多尺度嵌入解决高维数据可视化中的尺度难题。

The Tree-SNE Tree Exists

  • 引入额外尺度参数,将t-SNE从二维扩展为(2+1)维嵌入。
  • 证明在绝大多数初始条件下,最优嵌入随尺度参数连续变化。
  • 适用于多尺度聚类任务,适合需要层次化分析的科研人员。

高维数据的聚类与可视化是现代数据科学中的普遍任务。主流方法如t-SNE或UMAP面临‘尺度问题’:处理MNIST数据集时,应区分不同数字,还是区分同一数字的不同书写方式?答案取决于具体任务和尺度。本文重新审视Robinson & Pierce-Hoffman提出的观点,利用t-SNE中隐含的尺度对称性,将二维嵌入替换为(2+1)维嵌入,其中额外参数表示尺度。由此产生t-SNE树(简称tree-SNE)。我们证明,对于几乎所有初始条件,最优嵌入关于尺度参数连续依赖,即tree-SNE树存在。该思想可推广至其他吸引-排斥类方法,并在多个示例中得到验证。

原文摘要 · Abstract (English)

The clustering and visualisation of high-dimensional data is a ubiquitous task in modern data science. Popular techniques include nonlinear dimensionality reduction methods like t-SNE or UMAP. These methods face the `scale-problem' of clustering: when dealing with the MNIST dataset, do we want to distinguish different digits or do we want to distinguish different ways of writing the digits? The answer is task dependent and depends on scale. We revisit an idea of Robinson & Pierce-Hoffman that exploits an underlying scaling symmetry in t-SNE to replace 2-dimensional with (2+1)-dimensional embeddings where the additional parameter accounts for scale. This gives rise to the t-SNE tree (short: tree-SNE). We prove that the optimal embedding depends continuously on the scaling parameter for all initial conditions outside a set of measure 0: the tree-SNE tree exists. This idea conceivably extends to other attraction-repulsion methods and is illustrated on several examples.

可视化t-SNE多尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。