用构象集合建模提升环肽性质预测,效果优于单一结构方法。
Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

- 用共享等变图网络编码每个构象,再通过集合注意力聚合。
- 预训练后在数据集上达R²=0.477,比纯序列模型高0.038。
- 融合序列编码器可进一步提升性能,适合分子构象研究者。
从结构预测分子性质通常只用单一代表构象,但许多分子在溶液中以构象集合形式存在。本文提出EnsembleEGNN,一种基于循环肽构象集合的分子集合基础模型:先用共享的等变图神经网络(EGNN)编码每个构象,再通过集合注意力块聚合表示。在CREMP数据集上,采用多任务自监督目标(掩码词恢复、噪声坐标重建、成对距离重建)进行预训练。从零训练时,模型表现极差(R²=0.005)。而预训练模型在CREMP-CycPeptMPDB上达到R²=0.477,Pearson r=0.699,优于仅使用序列的BERT基线(R²=0.439,r=0.667)。当与BERT序列编码器端到端联合训练时,混合模型进一步提升至R²=0.538,r=0.737。结果表明,将热力学信息融入构象集合嵌入能有效提升环肽性质预测性能。
原文摘要 · Abstract (English)
Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representations with a Set Attention Block. We pretrain the model on CREMP, a cyclic peptide ensemble dataset, using a multi-task self-supervised objective combining masked token recovery, noisy-coordinate reconstruction, and pairwise distance reconstruction. On the CREMP-CycPeptMPDB dataset, training EnsembleEGNN from scratch fails entirely ($R^2=0.005$). However, the pretrained model reaches $R^2=0.477$ and Pearson $r=0.699$, outperforming the sequence-only BERT baseline ($R^2=0.439$, Pearson $r=0.667$). When EnsembleEGNN is co-trained end-to-end with the BERT sequence encoder, the hybrid model improves further to $R^2=0.538$ and Pearson $r=0.737$. These results demonstrate that encoding conformational ensembles into a single thermodynamically informed embedding improves cyclic-peptide property prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。