arXiv:2606.26872cs.CV2026-06

让图像编辑更精准:基于空间奖励的强化学习新方法

SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

论文配图:SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing
图 1 · 摘自论文原文
  • 引入区域感知奖励,将局部反馈转化为空间对齐的优化信号
  • 在多区域编辑任务中显著提升效果,优于传统方法
  • 适合需要精细控制图像局部内容的研究者

近期在线强化学习已显著提升图像编辑质量。然而,现有Flow-GRPO类方法通常依赖单一全图奖励,难以实现细粒度编辑优化。我们发现图像编辑的关键障碍在于空间均匀性假设:全图奖励无法区分不同空间区域对图像质量的贡献。为此,我们提出SpatialFlow-GRPO,一种引入空间细粒度奖励反馈的训练框架。该框架将区域感知奖励转换为语义区域级别的优化信号,并在策略更新中将区域优势与对应潜在位置对齐。我们还训练了一个区域感知奖励模型SFReward,构建了包含14,000个带区域标注的编辑样本的数据集SFReward-14K,并提出MultiEditBench用于评估多区域编辑能力。在OmniGen2和FLUX.2-klein-4B上,SpatialFlow-GRPO在GEdit-Bench、ImgEdit-Bench和MultiEditBench上均优于Flow-GRPO。结果表明,SpatialFlow-GRPO成功将局部反馈转化为空间对齐的更新信号,提升了编辑质量。

原文摘要 · Abstract (English)

Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which makes fine-grained editing optimization difficult. We observe that a key obstacle in image editing is this spatial uniformity assumption: a whole-image reward cannot distinguish how different spatial regions contribute to image quality. To address this issue, we propose SpatialFlow-GRPO, a training framework that introduces spatially fine-grained reward feedback. The framework converts region-aware rewards into semantic-region-level optimization signals and aligns region advantages with the corresponding latent positions during policy updates. We also train a region-aware reward model, SFReward, construct SFReward-14K with region-annotated editing samples, and introduce MultiEditBench to evaluate multi-region editing ability. On OmniGen2 and FLUX.2-klein-4B, SpatialFlow-GRPO outperforms Flow-GRPO on GEdit-Bench, ImgEdit-Bench, and MultiEditBench. The results show that SpatialFlow-GRPO converts local feedback into spatially aligned update signals and improves editing quality.

图像编辑强化学习空间注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。