用少量实验数据快速预测蛋白适应度,提升设计效率。
Few-shot Protein Fitness Prediction via In-context Learning and Test-time Training
- 基于上下文学习与测试时训练,动态适配新蛋白和检测任务。
- 在多种蛋白家族中表现优于零样本与全监督基线模型。
- 适合实验数据稀缺的蛋白工程研究者使用。
在蛋白工程中,仅用极少实验数据准确预测蛋白适应度始终是个挑战。本文提出PRIMO(PRotein In-context Mutation Oracle),一种基于Transformer的框架,利用上下文学习与测试时训练,无需大规模特定任务数据即可快速适应新蛋白和检测任务。通过将序列信息、辅助零样本预测及来自多个检测任务的稀疏实验标签统一编码为一个标记集,并在预训练的掩码语言建模范式下学习,PRIMO采用基于偏好的损失函数优先筛选有潜力的变异体。在涵盖多种蛋白家族及性质(包括替换与插入/缺失突变)的测试中,PRIMO均超越零样本与全监督基线模型。该工作凸显了大规模预训练与高效测试时适应相结合,在数据收集成本高、标签稀缺的蛋白设计任务中的强大潜力。
原文摘要 · Abstract (English)
Accurately predicting protein fitness with minimal experimental data is a persistent challenge in protein engineering. We introduce PRIMO (PRotein In-context Mutation Oracle), a transformer-based framework that leverages in-context learning and test-time training to adapt rapidly to new proteins and assays without large task-specific datasets. By encoding sequence information, auxiliary zero-shot predictions, and sparse experimental labels from many assays as a unified token set in a pre-training masked-language modeling paradigm, PRIMO learns to prioritize promising variants through a preference-based loss function. Across diverse protein families and properties-including both substitution and indel mutations-PRIMO outperforms zero-shot and fully supervised baselines. This work underscores the power of combining large-scale pre-training with efficient test-time adaptation to tackle challenging protein design tasks where data collection is expensive and label availability is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。