用元学习让蛋白语言模型在少数据下也能高效预测蛋白功能。
Metalic: Meta-Learning In-Context with Protein Language Models
- 通过元学习在多种蛋白功能任务中训练,实现跨任务快速适应。
- 在低数据场景下性能超越现有模型,参数量仅为其1/18。
- 适合蛋白工程、药物设计等缺乏标注数据的研究方向。
预测蛋白的生物物理与功能特性对计算机辅助蛋白设计至关重要。由于体外标注数据稀缺,现有模型往往缺乏目标任务的具体数据,通常依赖通用蛋白序列建模训练后进行微调或零样本预测。当无任务数据时,模型会假设蛋白序列概率与功能评分强相关,但这一假设常不成立。为此,我们提出Metalic(Meta-Learning In-Context),在标准蛋白功能预测任务分布上进行元学习,并利用上下文学习与微调适应新任务。关键在于,尽管元训练阶段未考虑微调,但微调后模型仍表现出显著泛化能力。我们的方法在参数量仅为当前最优模型1/18的情况下,实现了卓越性能;在ProteinGym基准测试中,于低数据设置下达到新纪录。鉴于数据稀缺,我们认为元学习将在推动蛋白工程发展中发挥关键作用。
原文摘要 · Abstract (English)
Predicting the biophysical and functional properties of proteins is essential for in silico protein design. Machine learning has emerged as a promising technique for such prediction tasks. However, the relative scarcity of in vitro annotations means that these models often have little, or no, specific data on the desired fitness prediction task. As a result of limited data, protein language models (PLMs) are typically trained on general protein sequence modeling tasks, and then fine-tuned, or applied zero-shot, to protein fitness prediction. When no task data is available, the models make strong assumptions about the correlation between the protein sequence likelihood and fitness scores. In contrast, we propose meta-learning over a distribution of standard fitness prediction tasks, and demonstrate positive transfer to unseen fitness prediction tasks. Our method, called Metalic (Meta-Learning In-Context), uses in-context learning and fine-tuning, when data is available, to adapt to new tasks. Crucially, fine-tuning enables considerable generalization, even though it is not accounted for during meta-training. Our fine-tuned models achieve strong results with 18 times fewer parameters than state-of-the-art models. Moreover, our method sets a new state-of-the-art in low-data settings on ProteinGym, an established fitness-prediction benchmark. Due to data scarcity, we believe meta-learning will play a pivotal role in advancing protein engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。