提出新采样方法,提升蛋白质语言模型设计抗体的效率。
How to make the most of your masked language model for protein engineering
- 用随机束搜索优化掩码语言模型生成序列。
- 实验证明采样方法对抗体设计效果影响显著。
- 适合从事蛋白质工程与生成模型研究者参考。
近年来涌现出大量蛋白质语言模型,但如何最优地从这些模型中采样以优化目标生物特性仍缺乏系统研究。本文提出一种灵活高效的掩码语言模型(MLM)采样方法,并在真实抗体治疗项目中进行了体内外系统的评估。首先,提出基于随机束搜索的采样策略,利用MLM高效评估单氨基酸替换全邻域伪困惑度的能力,将生成问题重构为整个序列的评估任务,从而支持多目标优化。其次,通过大规模体外对比实验发现,采样方法的选择对结果有显著影响,凸显了该领域研究的重要性。本工作填补了蛋白质生成模型应用中的关键空白。
原文摘要 · Abstract (English)
A plethora of protein language models have been released in recent years. Yet comparatively little work has addressed how to best sample from them to optimize desired biological properties. We fill this gap by proposing a flexible, effective sampling method for masked language models (MLMs), and by systematically evaluating models and methods both in silico and in vitro on actual antibody therapeutics campaigns. Firstly, we propose sampling with stochastic beam search, exploiting the fact that MLMs are remarkably efficient at evaluating the pseudo-perplexity of the entire 1-edit neighborhood of a sequence. Reframing generation in terms of entire-sequence evaluation enables flexible guidance with multiple optimization objectives. Secondly, we report results from our extensive in vitro head-to-head evaluation for the antibody engineering setting. This reveals that the choice of sampling method can have a substantial impact, motivating future research into this under-explored area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。