用三维原子信息预训练蛋白表示,提升结构预测精度
Atom-level Protein Representation Learning Improves Protein Structure Prediction

- 三视图联合建模氨基酸、主链与全原子几何信息
- 在多个结构预测任务中超越序列和已有结构表示模型
- 适合关注蛋白质结构建模与生成的研究者
生成模型的进步表明,预训练表示可作为生成的条件特征或对齐目标。受此启发,我们研究了超越传统功能注释的蛋白质结构预测中的表示学习。提出TriProRep,一种结构感知的预训练方法,通过离散编码的VQ-VAE分词器,联合建模三种对齐的残基级视图:氨基酸身份、主链几何和局部全原子几何。通过预训练恢复被生成器破坏的原始分词,TriProRep学会区分合理的但错误的跨视图增强与原始蛋白。进一步引入RepSP基准,用于评估表示在结构预测场景下的表现,测试三种用途:从无配体链表示预测同源二聚体共折叠、基于同源二聚体推断残基级相互作用属性、以及表示对齐的单体结构预测。在这些任务中,TriProRep优于仅使用序列和先前结构感知表示模型,同时在常规基准上保持竞争力。
原文摘要 · Abstract (English)
Recent advances in generative modeling show that pretrained representations can improve generation as conditioning features or alignment targets. Motivated by this, we study protein representations for predicting structures beyond conventional function annotation. We propose TriProRep, a structure-aware pretraining method that jointly models three aligned residue-level views: amino-acid identity, backbone geometry, and local full-atom geometry, discretely encoded via VQ-VAE tokenizers. By pretraining to recover original tokens from generator-corrupted views, TriProRep learns to distinguish plausible but incorrect cross-view augmentations from the original protein. We further introduce RepSP, a benchmark for evaluating protein representations in structure-predictive settings. RepSP tests three uses of representations: homodimer co-folding from apo-chain representations, residue-level prediction of homodimer-derived interaction properties, and representation-aligned monomer structure prediction. Across these tasks, TriProRep improves over sequence-only and prior structure-aware representation models, while maintaining competitive performance on conventional benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。