arXiv:2503.02039cs.LGcs.AI2025-03被引 40

提出动态搜索方法,让扩散模型生成时更精准匹配目标奖励。

Dynamic Search for Inference-Time Alignment in Diffusion Models

  • 将生成过程看作搜索问题,动态调整搜索范围和深度。
  • 在序列、分子和图像生成中,显著提升奖励优化效果。
  • 适合需要精确控制生成结果的科研与设计场景。

扩散模型在多个领域展现出强大的生成能力,但如何使其输出与期望的奖励函数对齐仍是挑战,尤其当奖励函数不可微时。已有无梯度引导方法难以实现最优推理阶段对齐。本文首次将推理阶段对齐建模为搜索问题,提出动态搜索扩散(DSearch)方法:通过从去噪过程中采样节点,近似中间节点的奖励,并动态调整束宽与树状扩展策略,高效探索高奖励生成路径。为优化中间决策,DSearch引入基于噪声水平的自适应调度机制和前瞻启发式函数。我们在生物序列设计、分子优化和图像生成等多个领域验证了该方法,结果表明其在奖励优化方面优于现有方法。

原文摘要 · Abstract (English)

Diffusion models have shown promising generative capabilities across diverse domains, yet aligning their outputs with desired reward functions remains a challenge, particularly in cases where reward functions are non-differentiable. Some gradient-free guidance methods have been developed, but they often struggle to achieve optimal inference-time alignment. In this work, we newly frame inference-time alignment in diffusion as a search problem and propose Dynamic Search for Diffusion (DSearch), which subsamples from denoising processes and approximates intermediate node rewards. It also dynamically adjusts beam width and tree expansion to efficiently explore high-reward generations. To refine intermediate decisions, DSearch incorporates adaptive scheduling based on noise levels and a lookahead heuristic function. We validate DSearch across multiple domains, including biological sequence design, molecular optimization, and image generation, demonstrating superior reward optimization compared to existing approaches.

扩散模型生成优化动态搜索奖励对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。