arXiv:2602.11216cs.LGphysics.bio-ph2026-02中稿 · ICML被引 4

用蛋白语言模型提升分子动力学迁移能力,更省数据且泛化更强。

Protein Language Model Embeddings Improve Generalization of Implicit Transfer Operators

  • 引入蛋白语言模型嵌入,增强隐式转移算子的数据效率
  • 在跨蛋白系统上实现当前最优的平衡采样性能
  • 适合需要高效泛化采样的生物分子模拟研究者

分子动力学(MD)是物理、化学和生物学中的核心计算工具,用于预测实验可观测量,其本质是对高维分子分布(如玻尔兹曼分布和转移密度)的期望计算。然而,传统MD受限于生成独立样本的高昂计算成本。生成式分子动力学(GenMD)作为替代方案,通过数据或与能量模型交互学习分子分布的代理模型。尽管采样效率高,但跨分子系统的迁移能力常受限。本文表明,引入辅助信息可提升可迁移隐式转移算子(TITO)的数据效率与泛化能力。我们发现粗粒度TITO模型显著优于玻尔兹曼模拟器,而引入蛋白语言模型(pLM)嵌入进一步提升了分布外泛化能力。所提方法PLaTITO在跨蛋白系统的平衡采样基准测试中表现最优,包括快速折叠蛋白。我们还研究了结构嵌入、温度及大语言模型衍生嵌入等额外条件信号对性能的影响。

原文摘要 · Abstract (English)

Molecular dynamics (MD) is a central computational tool in physics, chemistry, and biology, enabling quantitative prediction of experimental observables as expectations over high-dimensional molecular distributions such as Boltzmann distributions and transition densities. However, conventional MD is fundamentally limited by the high computational cost required to generate independent samples. Generative molecular dynamics (GenMD) has recently emerged as an alternative, learning surrogates of molecular distributions either from data or through interaction with energy models. While these methods enable efficient sampling, their transferability across molecular systems is often limited. In this work, we show that incorporating auxiliary sources of information can improve the data efficiency and generalization of transferable implicit transfer operators (TITO) for molecular dynamics. We find that coarse-grained TITO models are substantially more data-efficient than Boltzmann Emulators, and that incorporating protein language model (pLM) embeddings further improves out-of-distribution generalization. Our approach, PLaTITO, achieves state-of-the-art performance on equilibrium sampling benchmarks for out-of-distribution protein systems, including fast-folding proteins. We further study the impact of additional conditioning signals such as structural embeddings, temperature, and large-language-model-derived embeddings on model performance.

分子动力学蛋白语言模型迁移学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。