arXiv:2605.16348cs.LGcs.AI2026-05

让流模型高效利用反馈,生成结果可复用。

Flow-Direct: Feedback-Efficient and Reusable Guidance for Flow Models via Non-Parametric Guidance Field

论文配图:Flow-Direct: Feedback-Efficient and Reusable Guidance for Flow Models via Non-Parametric Guidance Field
图 1 · 摘自论文原文
  • 用累积样本构建持续更新的引导场,替代临时梯度。
  • 每条反馈都用于改进全局引导,提升反馈利用效率。
  • 优化后可直接复用引导场生成新样本,支持多目标协同。

无需训练的引导方法使预训练扩散与流模型能通过外部黑盒奖励函数优化特定目标。然而现有方法反馈效率低,因奖励信息仅临时用于局部梯度近似或离散搜索决策,随后即被丢弃。为此,我们提出Flow-Direct框架,通过一个持续存在的引导场指导生成过程。理论上,该引导场由基础分布与奖励加权目标分布之间的对数密度比解析导出,可将预训练分布迁移到目标分布。实践中,引导场以非参数估计器实现,基于所有累积的奖励评估样本构建。随着优化过程中样本增多,该经验引导场愈发精确。此持久化设计带来两大优势:第一,反馈高度高效——每次评估样本均用于优化全局引导场,无奖励信息浪费;第二,框架天然可复用——优化完成后,收集的数据集定义了可重复使用的引导场,可用于生成新目标样本而无需额外奖励评估,且不同引导场可组合生成满足多重目标的样本。

原文摘要 · Abstract (English)

Training-free guidance enables pre-trained diffusion and flow models to optimize application-specific objectives using feedback from external black-box reward functions. However, existing methods are feedback-inefficient because reward feedback is used only transiently to inform a localized gradient approximation or a discrete search decision, and is subsequently discarded. To address this limitation, we propose Flow-Direct, a framework that guides the generation process via a persistent guidance field. Theoretically, this guidance field is analytically derived from the log-density ratio between the base and reward-weighted target distributions; it transports the pre-trained distribution to the target distribution. In practice, the field is implemented as a non-parametric estimator constructed from all accumulated reward-evaluated samples. As more samples are collected during optimization, this empirical guidance field becomes increasingly accurate. This persistent formulation yields two major advantages. First, Flow-Direct is highly feedback-efficient: because every evaluated sample is used to refine the global guidance field, no reward information is wasted. Second, the framework is naturally reusable: once optimization is complete, the collected dataset defines a reusable guidance field for generating novel target samples without additional reward evaluations, and distinct guidance fields can be combined to generate samples that simultaneously satisfy multiple objectives.

流模型引导场反馈效率可复用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。