arXiv:2605.11347cs.LGcs.AI2026-05

无需反向传播的噪声优化框架,可对不可微生成模型进行奖励对齐。

Gradient-Free Noise Optimization for Reward Alignment in Generative Models

论文配图:Gradient-Free Noise Optimization for Reward Alignment in Generative Models
图 1 · 摘自论文原文
  • 提出零阶噪声优化(ZeNO),仅通过奖励评分更新噪声,无需梯度信息。
  • 在蛋白质结构生成任务中表现优异,即使无法反向传播也有效。
  • 适用于扩散模型、流模型等各类生成器,尤其适合不可微场景。

现有扩散模型和流模型的奖励对齐方法依赖多步随机轨迹,难以扩展到确定性生成器。噪声空间优化是自然替代方案,但现有方法需通过生成器和奖励路径进行反向传播,仅限于可微设置。为此,本文提出零阶噪声优化(ZeNO),将噪声优化建模为路径积分控制问题,仅需零阶奖励评估即可求解。当采用奥恩斯坦-乌伦贝克参考过程时,更新隐式对应朗之万动力学,目标为奖励加权分布。ZeNO支持推理时高效扩展,在多种生成器和奖励函数上表现强劲,包括在无法反向传播的蛋白质结构生成任务中也取得良好效果。

原文摘要 · Abstract (English)

Existing reward alignment methods for diffusion and flow models rely on multi-step stochastic trajectories, making them difficult to extend to deterministic generators. A natural alternative is noise-space optimization, but existing approaches require backpropagation through the generator and reward pipeline, limiting applicability to differentiable settings. To address this, here we present ZeNO (Zeroth-order Noise Optimization), a gradient-free framework that formulates noise optimization as a path-integral control problem, estimable from zeroth-order reward evaluations alone. When instantiated with an Ornstein--Uhlenbeck reference process, the update connects to Langevin dynamics implicitly targeting a reward-tilted distribution. ZeNO enables effective inference-time scaling and demonstrates strong performance across diverse generators and reward functions, including a protein structure generation task where backpropagation is infeasible.

生成模型奖励对齐零阶优化蛋白质生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。