给生成模型加记忆偏置,让蛋白质模拟更快发现新构象。
Learning Implicit Bias in Generative Spaces for Accelerating Protein Dynamics Emulation

- 在生成空间引入历史依赖的隐式偏置,引导采样避开旧结构。
- 在动态PDB数据集上多样性提升35%,速度最高快37倍。
- 适合需要高效探索稀有构象的蛋白质动力学研究者。
蛋白质动力学的生成模拟器能以极低成本生成合理轨迹,但其训练分布限制导致长期外推时易重复已知状态而难以触及稀有构象。受经典增强采样启发,我们为预训练模拟器的生成空间引入一种隐式、依赖历史的偏置。具体而言,一个感知历史的得分估计器,通过距离加权方式为冻结的模拟器添加偏置,使逆时间采样远离已生成结构,并由环境支持项正则化。为保持长期演化中的结构有效性,增加基于得分的精修步骤,利用冻结模拟器将漂移样本投影回数据流形。实验表明:(i) 在DynamicPDB-80上多样性提升35%;(ii) 在12个零样本快速折叠蛋白上,仅使用学习偏置即达到无偏模拟器覆盖度约15倍加速;结合精修后,加速达约37倍,且覆盖约3倍更多低能态。代码即将开源。
原文摘要 · Abstract (English)
Generative emulators of protein dynamics produce plausible trajectories at a fraction of the cost of molecular dynamics, but they inherit their training distribution and tend to revisit known states rather than reach rare ones under long-horizon extrapolation. Inspired by classical enhanced sampling, we introduce an implicit, history-dependent bias in the generative space of a pretrained emulator. Specifically, a history-aware score estimator augments the frozen emulator with a distance-weighted bias that steers reverse-time sampling away from previously generated structures, regularized by an environment-support term. To preserve structural validity at long horizons, a score-based refinement step re-projects drifted samples onto the data manifold using the frozen emulator. Our experiments demonstrate that the method (i) raises diversity by $35\%$ on DynamicPDB-80; (ii) on $12$ zero-shot Fast-Folding proteins, the learned bias alone reaches the unbiased emulator's coverage up to ${\sim}15\times$ faster, and pairing it with refinement reaches the coverage up to ${\sim}37\times$ faster while covering ${\sim}3\times$ as many low-energy states. Code will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。