通过空间推理提升图像编辑强化学习的评估精度
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
- 基于预测编辑区域进行像素级推理,实现精准评分
- 在260k数据集上训练,多基准测试达顶尖表现
- 适合需要高精度反馈的在线强化学习图像编辑场景
在线强化学习为复杂图像编辑提供了新路径,但受限于可靠且细粒度的奖励信号稀缺。现有评估器常出现‘注意力坍塌’问题,忽略跨图像比较,难以捕捉细节,导致感知不准与评分失衡。为此,我们提出SpatialReward,一种通过显式空间推理强制精确验证的奖励模型。该模型将推理锚定在预测的编辑区域,使语义判断基于像素级证据,显著提升评估准确性。在自建的260k空间感知数据集上训练后,其在MMRB2和EditReward-Bench上达到当前最优性能,并在新提出的MultiEditReward-Bench上超越专有评估器。此外,SpatialReward作为在线强化学习中的可靠信号,使OmniGen2在GEdit-Bench上提升+0.90,超过领先判别模型,且是GPT-4.1提升幅度(+0.45)的两倍。结果表明,空间推理对实现图像编辑的有效对齐至关重要。
原文摘要 · Abstract (English)
Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and fine-grained reward signals. Existing evaluators frequently struggle with a critical perception gap we term "Attention Collapse," where models neglect cross-image comparisons and fail to capture fine-grained details, resulting in inaccurate perception and miscalibrated scores. To address these limitations, we propose SpatialReward, a reward model that enforces precise verification via explicit spatial reasoning. By anchoring reasoning to predicted edit regions, SpatialReward grounds semantic judgments in pixel-level evidence, significantly enhancing evaluative accuracy. Trained on a curated 260k spatial-aware dataset, our model achieves state-of-the-art performance on MMRB2 and EditReward-Bench, and outperforms proprietary evaluators on our proposed MultiEditReward-Bench. Furthermore, SpatialReward serves as a robust signal in online RL, boosting OmniGen2 by +0.90 on GEdit-Bench--surpassing the leading discriminative model and doubling the gain of GPT-4.1 (+0.45). These results demonstrate that spatial reasoning is essential for unlocking effective alignment in image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。