用结构信息提升抗体优化效率,但语言模型软约束让纯序列方法也能追平。
Is Sequence Information All You Need for Bayesian Optimization of Antibodies?
- 引入蛋白质语言模型做软约束,引导搜索到高潜力区域
- 结构信息提升早期优化效率,但最终性能与序列方法相当
- 对稳定性优化,序列方法在加入软约束后可媲美结构方法
贝叶斯优化是抗体药物属性工程的理想选择,但其代理模型在高度结构化的抗体空间中难以确定,且可能因目标属性不同而异。目前尚无研究将结构信息纳入抗体贝叶斯优化。本文探索了多种融合结构信息的方法,并在结合亲和力和稳定性两个抗体属性的实验中,与多种仅依赖序列的方法进行对比。我们提出基于蛋白质语言模型的“软约束”机制,有效引导优化过程。结果表明,特定结构信息可提升稳定性优化的早期数据效率,但峰值性能与序列方法相当。当引入语言模型软约束后,亲和力优化的数据效率差距缩小,稳定性优化差距完全消除,序列方法性能与结构方法持平,引发对结构必要性的质疑。
原文摘要 · Abstract (English)
Bayesian optimization is a natural candidate for the engineering of antibody therapeutic properties, which is often iterative and expensive. However, finding the optimal choice of surrogate model for optimization over the highly structured antibody space is difficult, and may differ depending on the property being optimized. Moreover, to the best of our knowledge, no prior works have attempted to incorporate structural information into antibody Bayesian optimization. In this work, we explore different approaches to incorporating structural information into Bayesian optimization, and compare them to a variety of sequence-only approaches on two different antibody properties, binding affinity and stability. In addition, we propose the use of a protein language model-based ``soft constraint,'' which helps guide the optimization to promising regions of the space. We find that certain types of structural information improve data efficiency in early optimization rounds for stability, but have equivalent peak performance. Moreover, when incorporating the protein language model soft constraint we find that the data efficiency gap is diminished for affinity and eliminated for stability, resulting in sequence-only methods that match the performance of structure-based methods, raising questions about the necessity of structure in Bayesian optimization for antibodies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。