arXiv:2607.16263cs.LGcs.CE2026-07中稿 · ICML

用弱监督提升抗体表达量排序,解决实验数据少的难题

Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision

论文配图:Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision
图 1 · 摘自论文原文
  • 用偏好学习融合少量实验数据与大量免疫数据
  • 在1254个标注序列上表现优于基线模型
  • 适合抗体设计中数据稀缺场景的研究者

抗体表达量排序在抗体设计中至关重要,但受限于标注实验数据稀缺。为此,我们提出统一的基于偏好的学习框架,将稀疏的定量表达数据与来自免疫数据的大规模弱正样本监督相结合。通过引入联合掩码似然近似和IMGT对齐,我们将直接偏好优化(DPO)适配至蛋白质语言模型,实现对可变长度序列的高效训练。在包含1254个标注序列及400万未标注骆驼源抗体的内部数据集上评估显示,该方法在多数指标上均持续优于基线模型。结果表明,偏好学习能有效利用弱监督,在数据受限环境下提供可扩展的抗体表达优化方案。

原文摘要 · Abstract (English)

Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that integrates scarce quantitative expression data with large-scale weak positive supervision from immunization data. We adapt Direct Preference Optimization (DPO) to protein language models by introducing a union-masked log-likelihood approximation and IMGT-based alignment, enabling efficient training on variable-length sequences. Evaluating on a diverse internal dataset of 1254 labeled sequences and 4 million unlabeled camelid-derived antibodies, we show that our method consistently outperforms baselines on most metrics. Our results demonstrate that preference learning can effectively learn from weak supervision, providing a scalable solution for antibody expressibility optimization in data-constrained settings. Project page: https://kisoji-biotechnology-inc.github.io/Preference-Expression-Ranking/.

抗体设计偏好学习弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。