用元学习提升蛋白突变预测跨任务泛化能力
Meta-Learning for Cross-Task Generalization in Protein Mutation Property Prediction
- 采用MAML元学习框架,快速适应新任务
- 跨任务测试中功能适应性准确率高29%,训练时间少65%
- 适合数据少的工业级蛋白设计场景
蛋白突变对生物功能有深远影响,准确预测其性质变化对药物研发、蛋白工程和精准医疗至关重要。现有方法依赖为每个数据集微调特定Transformer模型,但在跨数据集泛化上受限于实验条件差异和目标域数据不足。本文提出两项创新:(1)首次将模型无关元学习(MAML)应用于蛋白突变性质预测;(2)引入分隔符标记的突变编码策略,直接将突变信息融入序列上下文。基于Transformer架构结合MAML,实现仅需少量梯度更新即可快速适应新任务,而非学习数据集特异性模式。该突变编码解决标准Transformer将突变位点视为未知标记导致性能下降的问题。在三个不同蛋白突变数据集(功能适应性、热稳定性、溶解度)上的评估显示显著优势:跨任务测试中,功能适应性准确率提升29%,训练时间减少65%;溶解度任务准确率提升94%,训练速度加快55%。该框架在不同数据规模下保持一致训练效率,特别适用于实验数据有限的工业应用和早期蛋白设计。本研究系统建立了元学习在蛋白突变分析中的应用范式,并提出高效突变编码方案,为蛋白工程中的跨域泛化提供变革性方法。
原文摘要 · Abstract (English)
Protein mutations can have profound effects on biological function, making accurate prediction of property changes critical for drug discovery, protein engineering, and precision medicine. Current approaches rely on fine-tuning protein-specific transformers for individual datasets, but struggle with cross-dataset generalization due to heterogeneous experimental conditions and limited target domain data. We introduce two key innovations: (1) the first application of Model-Agnostic Meta-Learning (MAML) to protein mutation property prediction, and (2) a novel mutation encoding strategy using separator tokens to directly incorporate mutations into sequence context. We build upon transformer architectures integrating them with MAML to enable rapid adaptation to new tasks through minimal gradient steps rather than learning dataset-specific patterns. Our mutation encoding addresses the critical limitation where standard transformers treat mutation positions as unknown tokens, significantly degrading performance. Evaluation across three diverse protein mutation datasets (functional fitness, thermal stability, and solubility) demonstrates significant advantages over traditional fine-tuning. In cross-task evaluation, our meta-learning approach achieves 29% better accuracy for functional fitness with 65% less training time, and 94% better accuracy for solubility with 55% faster training. The framework maintains consistent training efficiency regardless of dataset size, making it particularly valuable for industrial applications and early-stage protein design where experimental data is limited. This work establishes a systematic application of meta-learning to protein mutation analysis and introduces an effective mutation encoding strategy, offering transformative methodology for cross-domain generalization in protein engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。