arXiv:2605.18552cs.LGq-bio.BM2026-05中稿 · ICML被引 1

构建大规模非冗余蛋白折叠分类基准,提出高效自监督学习方法。

Protein Fold Classification at Scale: Benchmarking and Pretraining

论文配图:Protein Fold Classification at Scale: Benchmarking and Pretraining
图 1 · 摘自论文原文
  • 用高掩码率和几何不变编码器设计自监督模型,提升结构表征能力。
  • 在TEDBench上超越监督模型和现有基线,实现强性能且可扩展。
  • 适用于需要高效蛋白折叠分类的研究者,尤其关注结构生物学与深度学习交叉。

蛋白拓扑分类对解析生物功能至关重要,但受限于缺乏大规模、无重复的基准及难以扩展的模型。本文构建了基于域百科(TED)和Foldseek聚类的AlphaFold结构的大型非冗余基准TEDBench。实验表明,当前蛋白表示学习方法或需超大模型,或表现不佳。为此,我们提出掩码不变自编码器(MiAE),一种自监督框架:采用高达90%的掩码率,结合SE(3)不变编码器与轻量解码器,从潜在表示和掩码标记重建主链坐标。MiAE具有良好的可扩展性,在TEDBench上优于监督模型和先进基线,确立了蛋白折叠分类的有效范式。为验证迁移能力,我们在CATH v4.4的实验结构数据集上进一步评估。TEDBench已开源(https://github.com/BorgwardtLab/TEDBench)。

原文摘要 · Abstract (English)

Classifying protein topology is essential for deciphering biological function, but progress is held back by the lack of large-scale benchmarks that avoid duplicates and by models that do not scale well. We introduce TEDBench, a large-scale, non-redundant benchmark for protein fold classification constructed from the Encyclopedia of Domains (TED) and Foldseek-clustered AlphaFold structures. We show that on TEDBench, current protein representation learning methods either require very large models or fail to deliver strong performance. To address this challenge, we propose Masked Invariant Autoencoders (MiAE), a self-supervised framework for protein structure representation learning. MiAE uses an extremely high masking ratio of up to 90% with an $\mathrm{SE(3)}$-invariant encoder and a lightweight decoder that reconstructs backbone coordinates from the latent representation and mask tokens. MiAE scales well and outperforms supervised counterparts and state-of-the-art baselines on TEDBench, establishing a strong recipe for protein fold classification. To test transfer beyond AlphaFold structures, we further benchmark on a curated dataset from experimental structures of CATH v4.4. TEDBench is available at https://github.com/BorgwardtLab/TEDBench.

蛋白折叠自监督学习结构生物学表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。