让蛋白语言模型一键定制,精准预测特定蛋白结构与功能。
One protein is all you need
- 通过测试时训练,实时定制模型以适配单个目标蛋白。
- 在难预测蛋白上提升结构精度,实现蛋白适应性预测新纪录。
- 适合需要高精度个体蛋白分析的生物实验研究者使用。
机器学习在生物领域的泛化能力仍是核心挑战。现有自监督预训练方法虽能提升整体表现,却难以在特定蛋白上达到最优,而实验人员常需精准预测未在训练数据中的个别蛋白。为此,我们提出蛋白质测试时训练(ProteinTTT)方法,可在无需额外数据的前提下,对单个目标蛋白实时定制蛋白语言模型。该方法在不同模型、规模和数据集上均显著提升泛化性能:在困难蛋白上改善结构预测,实现蛋白适应性预测新纪录,并在两项功能预测任务中表现更优。两个案例研究显示,ProteinTTT在抗体-抗原环建模中更准确,使大奇幻病毒数据库中19%的结构得到优化,优于通用版AlphaFold2与ESMFold的表现。
原文摘要 · Abstract (English)
Generalization beyond training data remains a central challenge in machine learning for biology. A common way to enhance generalization is self-supervised pre-training on large datasets. However, aiming to perform well on all possible proteins can limit a model's capacity to excel on any specific one, whereas experimentalists typically need accurate predictions for individual proteins they study, often not covered in training data. To address this limitation, we propose a method that enables self-supervised customization of protein language models to one target protein at a time, on the fly, and without assuming any additional data. We show that our Protein Test-Time Training (ProteinTTT) method consistently enhances generalization across different models, their sizes, and datasets. ProteinTTT improves structure prediction for challenging targets, achieves new state-of-the-art results on protein fitness prediction, and enhances function prediction on two tasks. Through two challenging case studies, we also show that customization via ProteinTTT achieves more accurate antibody-antigen loop modeling and enhances 19% of structures in the Big Fantastic Virus Database, delivering improved predictions where general-purpose AlphaFold2 and ESMFold struggle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。