arXiv:2606.27939cs.LGcs.AI2026-06被引 1

用两阶段微调让蛋白序列精准匹配目标氨基酸组成。

Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition

  • 先领域微调,再用强化学习迭代优化氨基酸组成
  • 相比单阶段微调,能更精准满足特定氨基酸比例
  • 适合需要精确控制营养成分的合成蛋白设计

蛋白质语言模型是生物序列生成的标准先验,但如何引导其达到明确的分布目标仍缺乏研究。本文研究了一类约束性蛋白生成问题:序列需匹配目标氨基酸(AA)组成,同时保持合理的序列统计特性和多样性。核心应用是合成饲料蛋白设计,其中蛋白质的氨基酸组成直接影响营养价值。提出两阶段流程:首先在领域内数据集上进行领域自适应微调(FT),然后通过强化学习(RL)进行迭代奖励加权微调,以初始微调模型为固定参考。在两个氨基酸组成目标上评估,发现微调使平均组成接近目标,而后续强化学习则实现微调无法满足的特定序列约束。进一步对比了所提组成奖励项与两个基线及消融变体,分离了各训练阶段贡献,并验证了氨基酸组成对齐过程中未损害序列质量。

原文摘要 · Abstract (English)

Protein language models are standard priors for biological sequence generation, but steering them toward explicit distributional design targets remains largely unexplored. We study a constrained protein generation problem in which sequences must match a desired amino-acid (AA) composition profile while preserving plausible sequence statistics and diversity. The motivating application is synthetic feed protein design, where the AA composition of dietary proteins directly determines their nutritional value. We propose a two-stage pipeline in which domain-adaptive fine-tuning (FT) on an in-domain protein dataset is followed by iterative reward-weighted FT via reinforcement learning (RL) anchored against the FT model as a frozen reference. We evaluate the pipeline on two AA compositions and find that FT brings the average composition close to the target, while the subsequent RL enforces specific sequence constraints that FT alone cannot satisfy. We additionally evaluate the design choices of the proposed composition reward term against two baselines and an ablated variant, isolate the contribution of each training stage, and verify that AA composition alignment is achieved without degrading sequence quality.

蛋白生成氨基酸组成两阶段微调强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。