用知识蒸馏解决蛋白质设计中的多目标冲突与遗忘问题。
ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design

- 通过自蒸馏机制在学生模型上实现多教师偏好对齐。
- 在不损失设计可塑性的前提下,目标偏好提升8倍训练速度。
- 适合需要高效平衡多个设计目标的蛋白质工程研究者。
设计具有特定功能或特性的蛋白质是合成生物学和药物发现的核心目标。近年来,蛋白质语言模型(PLMs)实现了高度可设计的蛋白序列生成,而偏好对齐则为引导设计向期望功能和特性演进提供了有效途径。然而,现有方法常导致预训练知识的灾难性遗忘,损害基本设计能力,并难以平衡多个相互竞争的目标。为此,我们受基于策略的知识蒸馏(OPD)启发,提出ProteinOPD——一种多目标偏好对齐框架,能够在保持PLM固有设计能力的同时有效平衡多个偏好目标。ProteinOPD将预训练的PLM转化为特定偏好的教师模型,通过在学生模型自身轨迹上的令牌级OPD,将其知识蒸馏至共享学生模型。在此过程中,学生模型被对齐到加权教师的归一化几何共识,同时在目标冲突时保证优化边界。该方法弥合了OPD在多目标/多教师对齐中的空白。大量实验表明,ProteinOPD在不牺牲设计可塑性的前提下显著提升了目标偏好性能,相较基于强化学习的对齐方法实现8倍训练加速。
原文摘要 · Abstract (English)
Designing proteins with desired functions or properties represents a core goal in synthetic biology and drug discovery. Recent advances in protein language models (PLMs) have enabled the generation of highly designable protein sequences, while preference alignment provides a promising way to steer designs toward desired functions and properties. Nevertheless, they often trigger catastrophic forgetting of pretrained knowledge, degrading basic designability and failing to balance multiple competing objectives. To address these issues, we draw inspiration from On-Policy Distillation (OPD), an advanced post-training method renowned for mitigating catastrophic forgetting through its mode-seeking nature. In this work, we propose ProteinOPD, a multi-objective preference alignment framework that can effectively balance multiple preference objectives while maintaining the inherent designability of PLMs. ProteinOPD adapts a pretrained PLM into preference-specific teachers and distills their knowledge into a shared student via token-level OPD on the student's own trajectories. During this process, the student is aligned to a unique normalized geometric consensus of weighted teachers while ensuring bounded optimization under conflicts. This bridges the gap for OPD in multi-objective/teacher alignment. Extensive experiments show that ProteinOPD achieves substantial gains on target preference objectives without compromising the designability, with an 8x training speedup over RL-based alignment competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。