无需标注数据,让蛋白语言模型自我优化生成更可控的蛋白质。
Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization
- 用模型自身不确定性与语义一致性构建无监督奖励信号
- 两种新算法在多种温度下逼近最优性能,覆盖更广
- 适合缺乏实验反馈或标注数据的生物分子设计场景
蛋白语言模型(PLMs)已成为可控生物分子设计的强大工具,但其后训练通常依赖昂贵的湿实验验证或精心构建的偏好数据集。为突破这一监督瓶颈,我们提出无监督奖励优化框架,实现无需真实标签的可调控蛋白生成。核心洞察是:任务无关的奖励信号(结合模型内在不确定性与外在语义一致性)在不同基础模型和温度设置下,与可控性度量高度相关。基于此,我们提出两种离线算法:软奖励优化(SRO)与二值化奖励优化(BRO),有效最大化由这些代理奖励诱导的经典强化学习人类反馈目标。在组合式分布外提示上的大量实验表明,两种方法显著优于对比基线(DPO、KTO),并在多个采样温度、模型规模和蛋白家族中逼近理想性能。此外,经无监督奖励微调的PLM在pass@k评估中始终比基线模型具有更高覆盖率。该框架通过模型自生成经验实现自我改进,为缺乏标注偏好或实验反馈的场景提供了一条可扩展的可控生物分子设计路径。
原文摘要 · Abstract (English)
Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets. To overcome this supervision bottleneck, we introduce unsupervised reward optimization of PLMs, a comprehensive framework for steerable protein generation without ground-truth labels. Our key insight is that task-agnostic rewards, which combine intrinsic model uncertainty with extrinsic semantic consistency informed by protein representation models, exhibit strong correlation with controllability measures across base models and temperature regimes. Building upon this discovery, we propose two offline algorithms: Soft Reward Optimization (SRO) and Binarized Reward Optimization (BRO), which effectively maximize the classical RLHF objective induced by these proxy rewards. Extensive experiments on compositional out-of-distribution prompts demonstrate that both methods significantly outperform competitive baselines (DPO, KTO), while approaching oracle performance across multiple sampling temperatures, model scales and protein families. Moreover, PLMs fine-tuned with unsupervised rewards can achieve consistently higher coverage compared to their base model in pass@k evaluations. By enabling self-improvement of PLMs through their own generated experience, our framework provides a scalable pathway toward controllable biomolecular design in settings where labeled preferences or experimental feedback are scarce or unavailable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。