研究生物推理模型后训练阶段如何影响性能与泛化能力。
How Post-Training Shapes Biological Reasoning Models

- 分阶段后训练重塑模型泛化能力,各阶段作用不同。
- 监督微调提升域内表现但过早导致域外性能下降。
- 强化学习可恢复域外性能,适合追求泛化的研究者。
针对生物学的科学推理模型结合语言模型与多模态生物数据基础模型,通过后训练构建。然而,各训练阶段如何影响推理与泛化仍不清晰。我们系统研究了在基因组学、转录组学和蛋白质领域中,100多个模型在骨干网络、持续预训练(CPT)、监督微调(SFT)和强化学习(RL)等条件下,对域内(ID)与域外(OOD)性能的影响。结果表明,每个后训练阶段以独特方式重塑泛化能力,并非均匀增益:CPT通过对齐生物语言提升下游性能;SFT持续提升域内表现,但导致域外性能先升后降;在强SFT检查点上使用对齐奖励的强化学习,可提升域外性能并部分恢复泛化能力。说明生物推理性能并非随监督或算力单调提升,而是取决于训练阶段的组合策略。固定后训练预算下,短时SFT+大范围RL+阶段适应性不对称,能实现最佳的域内-域外权衡。
原文摘要 · Abstract (English)
Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, and proteins. These models are built through post-training, yet how each stage shapes reasoning and generalization remains poorly understood. We study when post-training improves performance and when it induces over-specialization. Across genomics, transcriptomics, and proteins, we train and evaluate more than 100 biological reasoning models under controlled variation in backbone, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL), measuring both in-domain (ID) and out-of-domain (OOD) performance. We find that each post-training stage reshapes generalization in a distinct way rather than contributing uniform gains. CPT improves downstream performance by aligning models with biological language. SFT consistently increases ID performance but causes OOD performance to peak early and decline as models fit the training distribution. RL, when applied to strong SFT checkpoints with aligned rewards, improves OOD performance and partially recovers generalization. These results show that biological reasoning does not improve monotonically with additional supervision or compute. Instead, performance depends on how training stages are composed. Under fixed post-training budgets, the strongest ID-OOD trade-off comes from brief SFT, larger RL allocations, and asymmetric adaptation capacity across stages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。