对比预训练蛋白嵌入效果,发现微调后序列级表示更优
Exploring the limits of pre-trained embeddings in machine-guided protein design: a case study on predicting AAV vector viability
- 用AAV衣壳做案例,测试多种预训练嵌入方法
- 微调后序列级嵌入预测准确率最高,优于氨基酸级嵌入
- 提示稀疏或局部突变数据需微调才能发挥模型潜力
蛋白质序列的有效表征是基于机器学习的蛋白质设计核心。然而,蛋白质生物工程对序列表示提出独特挑战:实验数据集中的突变数量少,且分布稀疏或高度集中于局部区域,限制了序列级表征提取功能信号的能力。尽管此类比较研究至关重要,但目前仍较匮乏。本研究以腺相关病毒(AAV)衣壳为案例,系统评估多种ProtBERT和ESM2嵌入变体作为序列表示的效果。结果表明,未微调时,氨基酸级嵌入在监督任务中表现优于序列级表示,而后者在无监督设置中更有效;但最佳性能仅在使用特定任务标签进行微调后实现,且序列级表示表现最优。此外,显著改变序列表示所需的突变程度远超典型生物工程研究范围,表明在突变稀疏或高度局部化的数据集中,微调不可或缺。
原文摘要 · Abstract (English)
Effective representations of protein sequences are widely recognized as a cornerstone of machine learning-based protein design. Yet, protein bioengineering poses unique challenges for sequence representation, as experimental datasets typically feature few mutations, which are either sparsely distributed across the entire sequence or densely concentrated within localized regions. This limits the ability of sequence-level representations to extract functionally meaningful signals. In addition, comprehensive comparative studies remain scarce, despite their crucial role in clarifying which representations best encode relevant information and ultimately support superior predictive performance. In this study, we systematically evaluate multiple ProtBERT and ESM2 embedding variants as sequence representations, using the adeno-associated virus capsid as a case study and prototypical example of bioengineering, where functional optimization is targeted through highly localized sequence variation within an otherwise large protein. Our results reveal that, prior to fine-tuning, amino acid-level embeddings outperform sequence-level representations in supervised predictive tasks, whereas the latter tend to be more effective in unsupervised settings. However, optimal performance is only achieved when embeddings are fine-tuned with task-specific labels, with sequence-level representations providing the best performance. Moreover, our findings indicate that the extent of sequence variation required to produce notable shifts in sequence representations exceeds what is typically explored in bioengineering studies, showing the need for fine-tuning in datasets characterized by sparse or highly localized mutations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。