arXiv:2602.02425cs.LGq-bio.QM2026-02

用预训练蛋白模型压缩嵌入,直接生成高适配性蛋白变体。

Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization

  • 将蛋白语言模型嵌入压缩到紧凑潜在空间,保留进化知识。
  • 在AAV和GFP设计任务上达到当前最优性能,无需预测器引导采样。
  • 合成数据自举可提升数据稀缺场景下的优化效果。

蛋白适配度优化面临组合空间巨大、高适配性变体极度稀疏的挑战。现有方法或性能不足,或依赖计算成本高昂的基于梯度的采样。我们提出CHASE框架,通过将预训练蛋白语言模型的演化知识压缩至紧凑潜在空间,利用无分类器引导的条件流匹配模型,实现无需预测器引导的直接生成高适配性变体。在AAV与GFP蛋白设计基准测试中,CHASE表现达到当前最优。此外,我们证明在数据受限场景下,通过合成数据自举可进一步提升性能。

原文摘要 · Abstract (English)

Protein fitness optimization is challenged by a vast combinatorial landscape where high-fitness variants are extremely sparse. Many current methods either underperform or require computationally expensive gradient-based sampling. We present CHASE, a framework that repurposes the evolutionary knowledge of pretrained protein language models by compressing their embeddings into a compact latent space. By training a conditional flow-matching model with classifier-free guidance, we enable the direct generation of high-fitness variants without predictor-based guidance during the ODE sampling steps. CHASE achieves state-of-the-art performance on AAV and GFP protein design benchmarks. Finally, we show that bootstrapping with synthetic data can further enhance performance in data-constrained settings.

蛋白设计扩散模型生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。