用双维度拓扑方法实现无比对基因组分类,精度更高且有理论保障。
$p$-adic Bi-Filtrations for Topological Machine Learning on Genomic Sequences

- 结合p进制距离与频次L1距离构建双滤波拓扑结构
- 在12个基因组数据集上优于4个基线方法,低样本下最高提升21个百分点
- 适合小样本基因组分类,尤其对序列层级结构敏感的任务
我们提出pVR,一种用于无比对基因组序列分类的拓扑机器学习框架,融合p进制数与拓扑数据分析。每条DNA序列通过两个互补轴编码:基于k-mer前缀的p进制距离,捕捉层次位置结构;以及基于k-mer频率的组合L1距离,捕捉局部序列内容。两个距离共同参数化一个双滤波维托里斯-赖普斯复形,其每条序列的拓扑摘要作为标准机器学习分类器的特征。我们建立了理论保证:对度量扰动的稳定性、对素数选择的不变性,并解释了单一边的p进制轴拓扑信息不足,而双滤波可恢复非平凡同调。在12个基因组基准测试(28至500条序列,3至7类)中,pVR在六个低样本数据集中的三个上超越四个现有无比对基线,最高提升达21个百分点;仅在SARS-CoV-2变异体数据集上表现较差,因点突变破坏了层次假设;所有方法在大数据集上均饱和。此外,相比5亿参数的核苷酸Transformer v2零样本冻结嵌入,在三个低样本基准上提升6.7至11.4个百分点。pVR代码已公开于https://github.com/MAHI-Group/pVR。
原文摘要 · Abstract (English)
We introduce pVR, a topological machine learning framework for alignment-free genomic sequence classification that combines $p$-adic numbers with topological data analysis. Each DNA sequence is encoded along two complementary axes: a $p$-adic distance on $k$-mer prefixes, which captures hierarchical positional structure, and a compositional $L_1$ distance on $k$-mer frequencies, which captures local sequence content. The two distances jointly parameterise a bi-filtered Vietoris--Rips complex, and per-sequence topological summaries from this bi-filtration serve as features for standard machine learning classifiers. We establish theoretical guarantees for the construction: stability under metric perturbations and invariance to the choice of prime, alongside a result that explains why a single $p$-adic axis is topologically uninformative and why the bi-filtration recovers nontrivial homology. On twelve genomic benchmarks ($28$ to $500$ sequences, $3$ to $7$ classes), pVR outperforms four established alignment-free baselines on three of six low-sample datasets, with gains of up to $21$ percentage points; it underperforms only on a SARS-CoV-2 variant benchmark whose point-mutation divergence violates the hierarchical assumption, and all methods saturate in the large-sample regime. pVR also outperforms zero-shot frozen embeddings from the 500M-parameter Nucleotide Transformer v2 by $6.7$ to $11.4$ percentage points on three low-sample benchmarks. The pVR codebase is publicly available at https://github.com/MAHI-Group/pVR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。