通过多奖励动态优化,实现精准图像编辑与生成。
RewardFlow: Generate Images by Optimizing What You Reward
- 基于多奖励朗之万动力学,无需反演即可控制生成过程。
- 在多个基准上达到最高编辑保真度与组合一致性。
- 支持语义对齐、对象一致性和人类偏好等多目标协同优化。
我们提出RewardFlow,一种无需反演的推理时框架,通过多奖励朗之万动力学引导预训练扩散模型和流匹配模型。该方法统一了语义对齐、感知保真度、局部定位、对象一致性及人类偏好等互补可微奖励,并引入基于可微VQA的奖励,通过图文推理提供细粒度语义监督。为协调异构目标,设计提示感知自适应策略,从指令中提取语义原语,推断编辑意图,并在采样过程中动态调节奖励权重与步长。在多个图像编辑与组合生成基准上,RewardFlow实现了最先进的编辑保真度与组合对齐效果。
原文摘要 · Abstract (English)
We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward Langevin dynamics. RewardFlow unifies complementary differentiable rewards for semantic alignment, perceptual fidelity, localized grounding, object consistency, and human preference, and further introduces a differentiable VQA-based reward that provides fine-grained semantic supervision through language-vision reasoning. To coordinate these heterogeneous objectives, we design a prompt-aware adaptive policy that extracts semantic primitives from the instruction, infers edit intent, and dynamically modulates reward weights and step sizes throughout sampling. Across several image editing and compositional generation benchmarks, RewardFlow delivers state-of-the-art edit fidelity and compositional alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。