arXiv:2509.19604cs.LG2025-09被引 1

用多模态模型预测抗体重设计成功率,减少实验浪费。

Improved Therapeutic Antibody Reformatting through Multimodal Machine Learning

  • 结合序列与结构信息的多模态机器学习框架
  • 在新抗体无数据场景下仍保持高预测准确率
  • 优于大语言模型,适合抗体工程研发人员

现代治疗性抗体设计常将多个功能域组合成复杂分子,虽可拓展疾病适应症并提升安全性,但面临功能与稳定性不确定、难以合成等工程挑战。为此,我们开发了一种机器学习框架,用于预测抗体从一种格式重设计为另一种格式的成功率。该框架融合抗体序列与结构上下文,并采用贴近真实部署场景的评估协议。在真实世界抗体重设计数据集上的实验发现,大型预训练蛋白质语言模型(PLMs)并未优于简单且领域定制的多模态表示。尤其在最困难的‘新抗体、无数据’测试场景中,最佳多模态模型仍达到高预测准确率,可有效筛选有潜力的候选分子,显著减少无效实验投入。

原文摘要 · Abstract (English)

Modern therapeutic antibody design often involves composing multi-part assemblages of individual functional domains, each of which may be derived from a different source or engineered independently. While these complex formats can expand disease applicability and improve safety, they present a significant engineering challenge: the function and stability of individual domains are not guaranteed in the novel format, and the entire molecule may no longer be synthesizable. To address these challenges, we develop a machine learning framework to predict "reformatting success" -- whether converting an antibody from one format to another will succeed or not. Our framework incorporates both antibody sequence and structural context, incorporating an evaluation protocol that reflects realistic deployment scenarios. In experiments on a real-world antibody reformatting dataset, we find the surprising result that large pretrained protein language models (PLMs) fail to outperform simple, domain-tailored, multimodal representations. This is particularly evident in the most difficult evaluation setting, where we test model generalization to a new starting antibody. In this challenging "new antibody, no data" scenario, our best multimodal model achieves high predictive accuracy, enabling prioritization of promising candidates and reducing wasted experimental effort.

抗体设计多模态学习药物研发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。