提出一套高效分子优化方法,用反馈迭代提升生成质量。
On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

- 结合采样选择、奖励塑造与模型去偏,动态调整生成方向。
- 在6个分子亲和力任务中,比基线提升显著,尤其在需大偏离时更优。
- 适合需要少样本反馈的药物分子设计场景。
分子优化通常从预训练生成模型出发,该模型捕获了有效分子结构的广泛先验。然而测试阶段的目标并非从中采样,而是利用有限的评估预算,将生成导向任务相关的高收益分子。本文研究离散扩散模型在此在线适应问题中的设计空间。每个在线轮次涉及多个决策:选择哪些候选分子评估、如何将奖励转化为模型更新、是否重用历史反馈、以及如何控制对先验的偏离程度。这些因素多被孤立研究,尚未明确它们在完整在线适应流程中是否互补、冗余或冲突。本文在六个小分子结合亲和力任务和三个蛋白适应度任务上进行受控实验,发现采样策略、奖励塑造和模型去偏能互补提升收益,尤其在小分子任务中效果显著。回放机制进一步稳定学习过程,有效性惩罚则确保小分子探索保持在合法分子流形上。综合来看,该方案提供了一套高效的反馈优化范式:在线微调结合采样选择、奖励塑造、去偏、回放与有效性控制。在相同评估预算和GPU小时数下,该方法优于离线微调与推理时搜索基线,当高收益分子需大幅偏离预训练先验时,优势最明显。
原文摘要 · Abstract (English)
Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures. At test time, however, the goal is not to sample from this prior, but to use a limited oracle budget to shift generation toward task-specific high-reward molecules. We study this adaptation problem for discrete diffusion models. Each online round couples several choices. The loop must decide which candidates to evaluate, how rewards become model updates, which feedback to reuse, and how far to move beyond the pretrained prior. These choices have mostly been studied in isolation, leaving open whether they complement one another, become redundant, or interfere inside a full online adaptation loop. We conduct controlled studies across six small-molecule binding-affinity tasks and three protein-fitness tasks. We find that acquisition, reward shaping, and model debiasing provide complementary routes to higher reward, especially for small molecules. Replay further stabilizes learning, while validity penalties keep small-molecule exploration on the valid molecular manifold. Together, these findings point to a practical recipe for feedback-efficient molecular optimization: online fine-tuning with acquisition, reward shaping, debiasing, replay, and validity control. This recipe outperforms offline fine-tuning and inference-time search baselines under matched oracle-call budgets and GPU-hour accounting. The gains are largest when high-reward candidates require larger shifts from the pretrained prior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。