arXiv:2503.09767cs.LGcs.CG2025-03ICML被引 3

学习可优化的拓扑覆盖,提升大规模数据的几何表示能力

Cover Learning for Large-Scale Topology Representation

  • 从优化角度学习数据覆盖,构建保拓扑的单纯复形
  • 生成的复形在规模和拓扑表达上优于传统方法
  • 适合需要高效大尺度拓扑建模的研究者

经典无监督学习方法如聚类和线性降维适用于离散或线性的大规模几何结构;现代流形学习通过构建图来寻找低维表示或推断局部几何。近期拓扑数据分析引入单纯复形表示数据拓扑,主要方法包括基于几何复形的拓扑推断与基于Mapper图的大规模拓扑可视化——核心是拓扑中的重叠构造(nerve construction),即由空间的子集覆盖构建单纯复形。然而这些方法存在局限:几何复形随数据量增长计算开销大,Mapper图调参困难且仅保留低维信息。本文提出将覆盖学习作为独立问题研究,并从优化视角出发,提出一种学习保拓扑覆盖的方法。所获单纯复形在规模上优于标准拓扑推断方法,在大尺度拓扑表示上超越Mapper类算法。

原文摘要 · Abstract (English)

Classical unsupervised learning methods like clustering and linear dimensionality reduction parametrize large-scale geometry when it is discrete or linear, while more modern methods from manifold learning find low dimensional representation or infer local geometry by constructing a graph on the input data. More recently, topological data analysis popularized the use of simplicial complexes to represent data topology with two main methodologies: topological inference with geometric complexes and large-scale topology visualization with Mapper graphs -- central to these is the nerve construction from topology, which builds a simplicial complex given a cover of a space by subsets. While successful, these have limitations: geometric complexes scale poorly with data size, and Mapper graphs can be hard to tune and only contain low dimensional information. In this paper, we propose to study the problem of learning covers in its own right, and from the perspective of optimization. We describe a method for learning topologically-faithful covers of geometric datasets, and show that the simplicial complexes thus obtained can outperform standard topological inference approaches in terms of size, and Mapper-type algorithms in terms of representation of large-scale topology.

拓扑数据分析覆盖学习单纯复形

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。