arXiv:2604.25209cs.LGcs.AI2026-04

提出可忠实保持数据拓扑结构的降维方法,显著优于传统方法。

DiRe-RAPIDS: Topology-faithful dimensionality reduction at scale

论文配图:DiRe-RAPIDS: Topology-faithful dimensionality reduction at scale
图 1 · 摘自论文原文
  • 基于已知同调结构的噪声流形设计拓扑保真度评测基准
  • 在72.3万论文嵌入上,拓扑保留能力是UMAP的3-4倍
  • 兼顾速度与精度,适合高维数据的拓扑分析任务

UMAP和t-SNE等降维方法虽广泛用于高维数据可视化,但其局部邻域目标会保留采样噪声并扭曲全局拓扑结构。我们发现标准局部度量会奖励这种噪声记忆:表现最佳的嵌入会产生数据中本不存在的环路和孤立岛。为此,我们构建了一个基于具有已知同调结构的噪声流形的拓扑保真度基准,用以调优DiRe。实验显示,经优化的DiRe配置在分类任务上达到或超过GPU加速版UMAP性能,且在压力测试中精确恢复了第一贝蒂数(first Betti numbers)。在72.3万arXiv论文嵌入数据集上,DiRe在相当的运行时间内保留了比UMAP多3至4倍的拓扑结构。

原文摘要 · Abstract (English)

Dimensionality reduction methods such as UMAP and t-SNE are central tools for visualising high-dimensional data, but their local-neighborhood objectives can preserve sampling noise while distorting global topology. We show that standard local metrics reward this noise memorisation: top-performing embeddings invent cycles and disconnected islands absent from the data. We introduce a topology-faithfulness benchmark based on noisy manifolds with known homology, tune DiRe against it, and find Pareto-optimal configurations that match or beat GPU-accelerated UMAP on classification while recovering exact first Betti numbers on stress tests. On 723K arXiv paper embeddings, DiRe preserves 3-4 times more topological structure than UMAP at comparable wall-clock.

降维拓扑保持高维数据UMAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。