arXiv:2504.16886q-bio.QMcs.LG2025-04被引 4

用预训练结构预测蛋白突变适应度,零样本即可评估功能影响。

Exploring zero-shot structure-based protein fitness prediction

  • 基于预测结构构建零样本蛋白适应度模型,无需额外标注数据。
  • 发现无序区域的结构预测会误导结果,降低预测性能。
  • 多模态集成模型表现强劲,适合作为基准对比方法。

利用预训练机器学习模型进行零样本蛋白序列变化的适应度预测,可广泛应用于遗传变异解读和蛋白质工程等下游任务,且无需额外标注数据。随着高精度蛋白结构预测工具的发展,大量预计算的结构数据得以生成,推动了基于结构的适应度预测模型的进步。我们通过实验评估了几种结构建模方式对下游适应度预测的影响。发现零样本模型在缺乏固定三维结构的无序区域上表现不佳;匹配实验中使用的结构与适应度检测条件至关重要,而这些区域的预测结构可能具有误导性,影响整体性能。最后,我们在ProteinGym替换基准上测试了一个新的结构基模型,结果表明简单的多模态集成已构成强有力的基线。

原文摘要 · Abstract (English)

The ability to make zero-shot predictions about the fitness consequences of protein sequence changes with pre-trained machine learning models enables many practical applications. Such models can be applied for downstream tasks like genetic variant interpretation and protein engineering without additional labeled data. The advent of capable protein structure prediction tools has led to the availability of orders of magnitude more precomputed predicted structures, giving rise to powerful structure-based fitness prediction models. Through our experiments, we assess several modeling choices for structure-based models and their effects on downstream fitness prediction. Zero-shot fitness prediction models can struggle to assess the fitness landscape within disordered regions of proteins, those that lack a fixed 3D structure. We confirm the importance of matching protein structures to fitness assays and find that predicted structures for disordered regions can be misleading and affect predictive performance. Lastly, we evaluate an additional structure-based model on the ProteinGym substitution benchmark and show that simple multi-modal ensembles are strong baselines.

蛋白设计零样本结构预测适应度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。