g-DPO让蛋白语言模型对齐更高效,训练速度提升1.7到5.4倍。
g-DPO: Scalable Preference Optimization for Protein Language Models
- 通过序列空间聚类剔除冗余训练对,保留关键信号。
- 用分组近似方法降低计算开销,实现更快收敛。
- 适合大规模蛋白设计数据,显著加速优化过程。
直接偏好优化(DPO)是将蛋白语言模型对齐实验设计目标的有效方法。然而,DPO存在可扩展性瓶颈:标注序列数增加时,训练对数量呈二次增长,导致即使是中小规模数据集也面临极长的训练时间。本文提出g-DPO框架,(i)利用序列空间聚类剔除冗余训练对,同时保留训练信号;(ii)通过分组近似方法分摊似然计算开销。在三个蛋白工程任务中,g-DPO的体外和体内性能与标准DPO无统计差异,但收敛速度提升1.7至5.4倍,且提速效果随数据集规模和突变景观结构增大而增强。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) is an effective approach for aligning protein language models with experimental design goals. However, DPO faces a scalability bottleneck: the number of possible training pairs grows quadratically with the number of labeled sequences, leading to prohibitive training times even for modestly sized datasets. We introduce g-DPO, a framework that (i) uses sequence space clustering to prune redundant pairs while preserving training signal, and (ii) amortizes likelihood computations with group-based approximations. Across three protein engineering tasks, g-DPO maintains in silico and in vitro performance that is statistically indistinguishable from standard DPO, while converging 1.7x to 5.4x times faster, with speedups that scale with dataset size and the structure of the underlying mutational landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。