arXiv:2605.03578q-bio.QMcond-mat.stat-mech2026-05被引 4

用高熵生成模型拓展人工蛋白序列空间,发现高熵模型更优。

Expanding functional protein sequence space using high entropy generative models

论文配图:Expanding functional protein sequence space using high entropy generative models
图 1 · 摘自论文原文
  • 采用渐进删减边的稀疏化方法构建高熵玻尔兹曼机模型
  • 高熵模型采样可行序列空间超低熵模型15个数量级
  • 适合追求广覆盖、抗过拟合的蛋白质设计研究者

基于进化序列数据训练的玻尔兹曼机已成为数据驱动设计人工蛋白的强大范式。然而,模型架构(尤其是参数密度)与实验性能之间的关系仍不明确。本文以分支酸变位酶家族为模型系统,比较了标准全连接玻尔兹曼机(bmDCA)与通过渐进边激活(eaDCA)和边删减(edDCA)生成的稀疏模型。在删减路径中识别出一个最大熵模型(meDCA),其在约束满足与概率分布灵活性间达到最优平衡。所有模型生成的人工序列均在体内互补实验中成功折叠为功能性酶,即使与天然序列差异显著也保持高成功率。尽管功能相当,meDCA模型所采样的可行序列空间比低熵模型大超过十五个数量级。进一步分析表明,高熵模型系统性降低过拟合,更好捕捉自然蛋白周围的局部中性空间。结果表明,虽多种满足共进化统计的模型均可生成功能性序列,但高熵玻尔兹曼机提供了对潜在进化适应度景观更优的表征。

原文摘要 · Abstract (English)

Boltzmann Machines trained on evolutionary sequence data have emerged as a powerful paradigm for the data-driven design of artificial proteins. However, the relationship between model architecture, specifically parameter density, and experimental performance remains poorly understood. Here, we investigate this relationship using the Chorismate Mutase enzyme family as a model system. We compare standard fully connected Boltzmann Machines for Direct Coupling Analysis (bmDCA) with sparse models generated via progressive edge activation (eaDCA) and edge decimation (edDCA). We identify a maximum-entropy model (meDCA) along the decimation trajectory that represents an optimal balance between constraint satisfaction and the flexibility of the probability distribution. We synthesized and tested artificial sequences from all models using an in vivo complementation assay, finding that all architectures, regardless of sparsity, generate functional enzymes with high success rates, even at significant divergence from natural sequences. Despite this functional equivalence, we demonstrate that the meDCA model samples a viable sequence space that is more than fifteen orders of magnitude larger than its low-entropy counterparts. Furthermore, comparative analyses reveal that high-entropy models systematically minimize overfitting and better capture the local neutral spaces surrounding natural proteins. These findings suggest that while various models satisfying coevolutionary statistics can generate functional sequences, high-entropy Boltzmann Machines provide a superior representation of the underlying evolutionary fitness landscape.

蛋白质设计生成模型高熵序列空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。