改进多模态蛋白模型结构建模,显著提升折叠精度与生成多样性
Elucidating the Design Space of Multimodal Protein Language Models
- 提出更精细的生成建模与结构感知架构,减少结构信息丢失
- 650M模型在PDB测试集上RMSD降至2.36,优于3B基线
- 适合蛋白质设计、结构预测及生成任务的研究者参考
多模态蛋白语言模型(PLMs)融合序列与基于标记的结构信息,为蛋白质建模、生成与设计提供强大基础。然而,依赖将三维结构离散化为标记会严重损失细粒度结构细节和相关性。本文系统揭示多模态PLMs的设计空间,识别出标记化损失与模型对结构标记预测不准为主要瓶颈。为此,我们提出涵盖改进生成建模、结构感知架构与表征学习、数据探索的新设计空间。进展实现更细粒度监督,证明基于标记的多模态PLMs可实现稳健的结构建模。有效设计方法显著提升结构生成多样性,并使650M模型在PDB测试集上将RMSD从5.52降低至2.36,甚至优于3B基线,接近专用折叠模型表现。
原文摘要 · Abstract (English)
Multimodal protein language models (PLMs) integrate sequence and token-based structural information, serving as a powerful foundation for protein modeling, generation, and design. However, the reliance on tokenizing 3D structures into discrete tokens causes substantial loss of fidelity about fine-grained structural details and correlations. In this paper, we systematically elucidate the design space of multimodal PLMs to overcome their limitations. We identify tokenization loss and inaccurate structure token predictions by the PLMs as major bottlenecks. To address these, our proposed design space covers improved generative modeling, structure-aware architectures and representation learning, and data exploration. Our advancements approach finer-grained supervision, demonstrating that token-based multimodal PLMs can achieve robust structural modeling. The effective design methods dramatically improve the structure generation diversity, and notably, folding abilities of our 650M model by reducing the RMSD from 5.52 to 2.36 on PDB testset, even outperforming 3B baselines and on par with the specialized folding models. Project page and code: https://bytedance.github.io/dplm/dplm-2.1/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。