arXiv:2511.13244cs.LGcs.AI2025-11

用遗传算法让蛋白质生成模型融入实验数据,突破不可导限制。

Seek and You Shall Fold

  • 用定制遗传算法连接扩散生成器与任意黑箱目标函数
  • 首次实现化学化学位移引导的蛋白质结构生成,准确率提升显著
  • 适合需要融合多源实验数据的结构生物学研究者

精确的蛋白质结构对理解生物功能至关重要,但将实验数据融入蛋白质生成模型仍是重大挑战。多数实验可观测值预测器不可微,无法与基于梯度的条件采样兼容,尤其在核磁共振领域,如化学位移等丰富数据难以直接整合到生成建模中。本文提出一种针对非可微信号的蛋白质生成模型引导框架,将连续扩散生成器与任意黑箱目标函数通过定制遗传算法耦合。我们在三种模态上验证其有效性:成对距离约束、核奥弗豪泽效应限制,以及首次实现的化学化学位移引导。结果表明化学化学位移引导结构生成可行,揭示了当前预测器的关键缺陷,并展示了一种通用的多样化实验信号整合策略。本工作为超越可微性限制的自动化、数据驱动蛋白质建模指明方向。

原文摘要 · Abstract (English)

Accurate protein structures are essential for understanding biological function, yet incorporating experimental data into protein generative models remains a major challenge. Most predictors of experimental observables are non-differentiable, making them incompatible with gradient-based conditional sampling. This is especially limiting in nuclear magnetic resonance, where rich data such as chemical shifts are hard to directly integrate into generative modeling. We introduce a framework for non-differentiable guidance of protein generative models, coupling a continuous diffusion-based generator with any black-box objective via a tailored genetic algorithm. We demonstrate its effectiveness across three modalities: pairwise distance constraints, nuclear Overhauser effect restraints, and for the first time chemical shifts. These results establish chemical shift guided structure generation as feasible, expose key weaknesses in current predictors, and showcase a general strategy for incorporating diverse experimental signals. Our work points toward automated, data-conditioned protein modeling beyond the limits of differentiability.

蛋白质结构生成模型实验数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。