通过噪声空间对齐实现分子属性可控生成,无需奖励模型即可调节生成结果。
Reward Transport: Property Control in Flow Matching via Noise-Space Alignment

- 利用最优传输对齐噪声与数据,将属性值映射到噪声坐标上
- 单一标量变量可单调控制logP和一致调控QED,且响应方向因目标而异
- 适用于无监督生成场景,与分类器自由引导等方法互补
流匹配中的噪声-数据配对通常被视为计算选择,我们发现其可作为对齐接口:通过根据目标分子属性匹配噪声与数据,能直接将可控结构嵌入学习到的流场中。基于此,我们提出Reward Transport,训练时使用最优传输将标量噪声空间坐标与分子奖励对齐;推理时仅通过调整该坐标即可引导生成分布,无需奖励模型、梯度指导或额外计算。在保持配对不变的极限下,阈值化该坐标可恢复交叉熵方法的截断奖励分布,提供一个连续可调的分布级控制旋钮。实验表明,在ZINC-250K和GuacaMol数据集上,该标量可单调控制logP,并在操作范围内一致调控QED;最显著的是,同一控制变量对不同目标产生相反结构响应——logP目标增长分子,QED目标缩小分子,排除了通用尺寸偏差的可能性。该接口与分类器自由引导和条件流匹配互补,而ε预测扩散中的负结果揭示了耦合级对齐在结构上的缺失。
原文摘要 · Abstract (English)
The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data according to a target molecular property, it embeds controllable structure directly into the learned flow field. Building on this view, we introduce Reward Transport, which uses optimal transport coupling at training time to align a scalar noise-space coordinate with molecular rewards; at inference, varying this coordinate steers the generated distribution without requiring an oracle, reward model, gradient guidance, or additional computation. In the coupling-preserving limit, thresholding this coordinate recovers the Cross-Entropy Method's truncated reward distribution, providing a principled, continuously adjustable distribution-level control knob. Empirically, on ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control over its operating range; most tellingly, the same knob produces opposite structural responses for different targets, growing molecules for logP but shrinking them for QED, which rules out a generic size bias. The interface is complementary to classifier-free guidance and conditional flow matching, while a negative result under epsilon-prediction diffusion clarifies where coupling-level alignment is structurally absent. Code: https://github.com/KehanGuo2/reward-transport
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。