将拓扑特征的持续性视为生存分析,统一实现假设检验与向量化。
From Persistence to Survival: Hypothesis Testing, Effect Sizes and Vectorisation for Topological Features

- 把持久性图看作生存数据,用生存函数统一建模
- 小样本下检验功效高,且类型I误差校准良好
- 适用于神经科学、图和点云等多类数据的下游任务
持久性图是拓扑数据分析中的常见表示,但不自然属于向量空间,其比较方法与下游预测工具长期分离。我们提出STRAND(生存拓扑图表示分析),将(一组)持久性图视为生存数据:每个拓扑特征的持久性值 $p = d - b$ 被视为完全观测的事件发生时间,持久性生存函数 $S(t) = \mathbb{P}(p > t)$ 成为比较图的核心对象。基于此单一表示,我们推导出:(i) 小样本下校准类型I误差且功效高的非参数两样本检验;(ii) 可解释的效果量;(iii) 1-Wasserstein稳定的特征向量用于下游机器学习。我们在具有可控拓扑结构的合成流形上验证了校准性和功效,在14个图和3个3D点云基准上展示出竞争力的向量化性能,并将其应用于功能性脑连接的fMRI/神经科学数据研究。据我们所知,STRAND是首个从单一一致且可解释表示中同时提供假设检验与向量化的持久性图方法。
原文摘要 · Abstract (English)
Persistence diagrams are common representations in topological data analysis, but they do not naturally live in a vector space, and the statistical tools developed for comparing them have largely evolved separately from those used for downstream prediction. We introduce STRAND (Survival Topological Representation ANalysis of Diagrams), which treats (collections of) PDs as survival data: each topological feature with persistence value $p = d - b$ is a fully observed time-to-event, and the persistence survival function $S(t) = \mathbb{P}(p > t)$ is the central object for comparing diagrams. From this single representation we derive (i) a non-parametric two-sample test with calibrated Type I error and high power from a small number of diagrams; (ii) interpretable effect sizes; and (iii) a 1-Wasserstein-stable feature vector for downstream machine learning. We validate calibration and power on synthetic manifolds with controlled topology, demonstrate competitive vectorisation across 14 graph and 3D point cloud benchmarks, and apply the method to study functional brain connectivity in fMRI/neuroscience data. To our knowledge, STRAND is the first method to provide hypothesis testing and vectorisation for persistence diagrams from a single coherent and interpretable representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。