提出区域约束策略优化方法,提升图像编辑的精准与保真度。
Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing
- 通过区域解耦噪声扰动,限制探索范围以减少干扰
- 在CompBench上实现编辑区域遵循度与非目标区域保留率双提升
- 适合需要高精度局部编辑的应用场景
指令引导的图像编辑需在目标修改与非目标保留之间取得平衡。近年来,基于流的模型因其高保真度和高效的确定性ODE采样,成为主流骨干架构。在此基础上,基于GRPO的奖励驱动后训练被用于直接优化编辑特定奖励,提升了指令遵循能力与编辑一致性。然而,现有方法常面临噪声信用分配问题:全局探索同时扰动非目标区域,导致组内奖励方差增大,使GRPO优势估计嘈杂。为此,我们提出RC-GRPO-Editing,一种基于确定性ODE采样的区域约束GRPO后训练框架。该方法抑制背景引入的冗余方差,实现更清晰的局部信用分配,在提升编辑区域指令遵循度的同时保持非目标内容不变。具体而言,通过区域解耦初始噪声扰动实现探索定位,降低背景引起的奖励方差并稳定GRPO优势;引入注意力聚焦奖励,使跨注意力在整个滚动过程中对齐预期编辑区域,减少非目标区域的意外变化。在CompBench上的实验表明,该方法在编辑区域指令遵循度与非目标保留方面均有持续改进。
原文摘要 · Abstract (English)
Instruction-guided image editing requires balancing target modification with non-target preservation. Recently, flow-based models have emerged as a strong and increasingly adopted backbone for instruction-guided image editing, thanks to their high fidelity and efficient deterministic ODE sampling. Building on this foundation, GRPO-based reward-driven post-training has been explored to directly optimize editing-specific rewards, improving instruction following and editing consistency. However, existing methods often suffer from noisy credit assignment: global exploration also perturbs non-target regions, inflating within-group reward variance and yielding noisy GRPO advantages. To address this, we propose RC-GRPO-Editing, a region-constrained GRPO post-training framework for flow-based image editing under deterministic ODE sampling. It suppresses background-induced nuisance variance to enable cleaner localized credit assignment, improving editing region instruction adherence while preserving non-target content. Concretely, we localize exploration via region-decoupled initial noise perturbations to reduce background-induced reward variance and stabilize GRPO advantages, and introduce an attention concentration reward that aligns cross-attention with the intended editing region throughout the rollout, reducing unintended changes in non-target regions. Experiments on CompBench show consistent improvements in editing region instruction adherence and non-target preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。