arXiv:2603.03375stat.MLcs.LG2026-03

修复UMAP理论漏洞,厘清其数学基础与算法对应关系。

The Theory behind UMAP?

  • 基于斯皮瓦克未发表手稿重构度量实现的完整推导
  • 揭示原始文献中的多处错误及其对后续研究的影响
  • 为数据科学家提供可信赖的理论依据

2018年,McInnes等人提出了广受数据科学界欢迎的降维算法UMAP。该工作基于斯皮瓦克未发表的手稿中一个名为度量实现的函子的有限变体。然而,该手稿存在诸多错误,这些错误被McInnes等人及后续研究沿用。本文旨在修正这些错误,提供斯皮瓦克函子与McInnes等有限变体的完整自包含推导。我们贡献了度量实现及相关函子的显式描述,并最终讨论了UMAP算法本身,以及对其性质和有限变体与原始算法对应关系的若干主张。

原文摘要 · Abstract (English)

In 2018, McInnes et al. introduced a dimensionality reduction algorithm called UMAP, which enjoys wide popularity among data scientists. Their work introduces a finite variant of a functor called the metric realization, based on an unpublished draft by Spivak. This draft contains many errors, most of which are reproduced by McInnes et al. and subsequent publications. This article aims to repair these errors and provide a self-contained document with the full derivation of Spivak's functors and McInnes et al.'s finite variant. We contribute an explicit description of the metric realization and related functors. At the end, we discuss the UMAP algorithm, as well as claims about properties of the algorithm and the correspondence of McInnes et al.'s finite variant to the UMAP algorithm.

降维理论分析拓扑学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。