arXiv:2505.15093q-bio.BMcs.LG2025-05NeurIPS被引 11

用少量实验数据指导生成模型,高效优化蛋白质功能。

Steering Generative Models with Experimental Data for Protein Fitness Optimization

  • 用数百个序列-功能配对数据,通过分类器引导和后验采样指导蛋白生成。
  • 在低通量实验条件下,生成模型性能优于强化学习方法。
  • 可直接嵌入自适应选择策略,适合实际生物实验优化场景。

蛋白质功能优化需在组合爆炸的序列空间中寻找最优序列。近期研究利用带标签数据引导生成模型(如扩散模型、语言模型)展现出潜力,但多数方法依赖代理奖励或大量标注数据,难以评估其在真实实验中表现。本研究聚焦仅几百个标签数据的情况,系统评估了分类器引导与后验采样在不同离散扩散模型上的效果,并展示如何将引导机制融入类似贝叶斯优化中的汤普森采样策略。结果表明,即插即用的引导方法优于基于蛋白语言模型的强化学习方案。研究为下一代蛋白质功能优化提供了实用指导。

原文摘要 · Abstract (English)

Protein fitness optimization involves finding a protein sequence that maximizes desired quantitative properties in a combinatorially large design space of possible sequences. Recent advances in steering protein generative models (e.g., diffusion models and language models) with labeled data offer a promising approach. However, most previous studies have optimized surrogate rewards and/or utilized large amounts of labeled data for steering, making it unclear how well existing methods perform and compare to each other in real-world optimization campaigns where fitness is measured through low-throughput wet-lab assays. In this study, we explore fitness optimization using small amounts (hundreds) of labeled sequence-fitness pairs and comprehensively evaluate strategies such as classifier guidance and posterior sampling for guiding generation from different discrete diffusion models of protein sequences. We also demonstrate how guidance can be integrated into adaptive sequence selection akin to Thompson sampling in Bayesian optimization, showing that plug-and-play guidance strategies offer advantages over alternatives such as reinforcement learning with protein language models. Overall, we provide practical insights into how to effectively steer modern generative models for next-generation protein fitness optimization.

蛋白质设计生成模型扩散模型小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。