用自蒸馏微调提升蛋白质语言模型设计能力,无需昂贵实验数据
Self Distillation Fine-Tuning of Protein Language Models Improves Versatility in Protein Design
- 利用模型自身生成数据,结合领域过滤器构建高质量训练集
- 微调后蛋白序列更稳定、功能更强,且多样性显著提升
- 方法通用性强,适合各类蛋白质模型与设计任务
监督微调(SFT)是将大模型适配到特定领域的标准方法,但其在蛋白质序列建模和蛋白质语言模型(PLM)中的应用仍缺乏系统性。主要原因在于蛋白质的高质量标注数据远难于自然语言获取。本文提出一种简单通用的快速SFT方案,旨在提升生成蛋白序列的保真度、可靠性和新颖性。与依赖昂贵预编实验数据的现有方法不同,本方法利用PLM自身,结合轻量级数据清洗流程与领域特异性过滤器,构建高质量训练数据。这些过滤器可独立优化PLM输出并筛选体外评估候选;与SFT结合后,使PLM能生成更稳定、功能更强的酶,同时拓展对非天然变体的蛋白序列空间探索。尽管该方法不依赖具体PLM或蛋白体系,我们以基因组规模的PLM(GenSLM)应用于色氨酸合酶家族,验证其有效性:微调后的模型生成序列不仅更具新颖性,且在目标设计约束和涌现蛋白属性指标上均有提升。
原文摘要 · Abstract (English)
Supervised fine-tuning (SFT) is a standard approach for adapting large language models to specialized domains, yet its application to protein sequence modeling and protein language models (PLMs) remains ad hoc. This is in part because high-quality annotated data are far more difficult to obtain for proteins than for natural language. We present a simple and general recipe for fast SFT of PLMs, designed to improve the fidelity, reliability, and novelty of generated protein sequences. Unlike existing approaches that require costly precompiled experimental datasets for SFT, our method leverages the PLM itself, integrating a lightweight curation pipeline with domain-specific filters to construct high-quality training data. These filters can independently refine a PLM's output and identify candidates for in vitro evaluation; when combined with SFT, they enable PLMs to generate more stable and functional enzymes, while expanding exploration into protein sequence space beyond natural variants. Although our approach is agnostic to both the choice of protein language model (PLM) and the protein system, we demonstrate its effectiveness with a genome-scale PLM (GenSLM) applied to the tryptophan synthase enzyme family. The supervised fine-tuned model generates sequences that are not only more novel but also display improved characteristics across both targeted design constraints and emergent protein property measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。