用未约化的边界矩阵生成拓扑特征,提升机器学习性能并降低计算开销。
Unreduced Persistence Diagrams for Topological Machine Learning
- 直接从未约化边界矩阵提取拓扑特征向量,保留更多原始信息。
- 在多个任务上表现优于或媲美传统约化图谱的模型,部分任务提升显著。
- 计算更高效:内存消耗低十倍,且可并行处理,适合大规模数据。
基于持久同调特征的监督学习流程在实验中被发现忽略了持久图谱中的大量信息。而计算持久图谱通常是此类流程中最耗时的步骤。为此,本文提出若干方法,从未约化边界矩阵生成拓扑特征向量,并研究其理论与计算特性。我们在多种数据集和任务类型上比较了基于未约化图谱向量化与完全约化图谱向量化模型的性能。结果表明,在某些任务上,使用未约化图谱构建的模型表现相当甚至更优。我们还对一种实现未约化图谱计算的算法进行了基准测试,该算法为Ripser的深度修改版本。结果显示,该计算过程可并行化,平均内存消耗比完整持久图谱低一个数量级。结果表明,利用未约化边界矩阵中的信息,拓扑特征驱动的机器学习流程可在计算成本和性能上获益。
原文摘要 · Abstract (English)
Supervised machine learning pipelines trained on features derived from persistent homology have been experimentally observed to ignore much of the information contained in a persistence diagram. Computing persistence diagrams is often the most computationally demanding step in such a pipeline, however. To explore this dynamic, we introduce several methods to generate topological feature vectors from unreduced boundary matrices and investigate their theoretical and computational properties. We compared the performance of pipelines trained on vectorizations of unreduced PDs to vectorizations of fully-reduced PDs across several data and task types. Our results indicate that models trained on PDs built from unreduced diagrams can perform on par and even outperform those trained on fully-reduced diagrams on some tasks. We also benchmarked the computational performance of an algorithm for computing unreduced diagrams, which was implemented as a heavily modified version of Ripser. These computations are parallelizable and required an order of magnitude less memory on average compared to computing full persistence diagrams. Our results suggest that machine learning pipelines which incorporate topology-based features may benefit in terms of computational cost and performance by utilizing information contained in unreduced boundary matrices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。