arXiv:2607.16553cs.LG2026-07中稿 · IEEE International…

用离散曲率做蛋白折叠分类,轻量高效且可解释。

Discrete Ricci Curvature on Protein Contact Graphs for Lightweight Fold Classification

论文配图:Discrete Ricci Curvature on Protein Contact Graphs for Lightweight Fold Classification
图 1 · 摘自论文原文
  • 用原子接触图的离散曲率生成22维结构特征。
  • 仅22维就超越平均池化的ESM-2模型,性能差距在40%数据集上更大。
  • 结合持久同调后表现最佳,适合追求效率与可解释性的研究者。

蛋白折叠分类通常依赖序列表示或结构描述符,但轻量级手工设计描述符与预训练蛋白质语言模型嵌入之间的直接比较仍较有限。本文将离散Ricci曲率应用于Cα接触图,作为轻量级结构描述符用于折叠分类。每个蛋白域通过22维固定长度特征表示,该特征基于Ollivier-Ricci与Forman-Ricci边曲率分布的统计量和分位数生成。在CATH top-10拓扑分类和ASTRAL 40%身份阈值下的SCOPe top-10折叠基准上进行评估,对比几何特征、接触图统计、持久同调及均值池化的ESM-2(150M)基线。在两个数据集上,轻量结构描述符显著优于均值池化的ESM-2嵌入,尤其在ASTRAL 40% SCOPe上差距更明显。仅使用22维的Ricci特征(仅为ESM-2的3.4%)即在两个数据集上超越其表现。结合持久同调后达到最优,特征向量为112维,在CATH上获得0.71宏F1,SCOPe上达0.68。结果表明,轻量可解释的图描述符可成为预训练模型嵌入的实用替代方案。

原文摘要 · Abstract (English)

Protein fold classification can be approached via sequence-based representations or structural descriptors, but direct comparisons between lightweight handcrafted descriptors and pretrained protein language model embeddings remain limited. We investigate discrete Ricci curvature on Calpha contact graphs as a lightweight structural descriptor for fold classification. Each protein domain is represented by a 22-dimensional fixed-length feature derived from summary statistics and quantiles of Ollivier-Ricci and Forman-Ricci edge curvature distributions. We evaluate on CATH top-10 Topology classification and on the ASTRAL 40%-identity SCOPe top-10 Fold benchmark, comparing against geometry, contact-graph statistics, persistent homology, and mean-pooled ESM-2 (150M) baselines. On both datasets, lightweight structural descriptors substantially outperform mean-pooled ESM-2 embeddings, with a larger performance gap on the ASTRAL 40% SCOPe benchmark. Ricci alone uses 22 dimensions, or 3.4% of the ESM-2 baseline dimensionality, and already outperforms mean-pooled ESM-2 on both datasets. Combining Ricci with persistent homology yields the strongest performance, achieving macro-F1 of 0.71 on CATH and 0.68 on SCOPe with a 112-dimensional feature vector. These results identify a regime where lightweight interpretable graph descriptors offer a practical alternative to pretrained protein language model embeddings.

蛋白折叠图神经网络曲率分析轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。