arXiv:2602.07458cs.CV2026-02中稿 · ICML被引 8

通过空间推理提升图像编辑强化学习的评估精度

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

  • 基于预测编辑区域进行像素级推理,实现精准评分
  • 在260k数据集上训练,多基准测试达顶尖表现
  • 适合需要高精度反馈的在线强化学习图像编辑场景

在线强化学习为复杂图像编辑提供了新路径,但受限于可靠且细粒度的奖励信号稀缺。现有评估器常出现‘注意力坍塌’问题,忽略跨图像比较,难以捕捉细节,导致感知不准与评分失衡。为此,我们提出SpatialReward,一种通过显式空间推理强制精确验证的奖励模型。该模型将推理锚定在预测的编辑区域,使语义判断基于像素级证据,显著提升评估准确性。在自建的260k空间感知数据集上训练后,其在MMRB2和EditReward-Bench上达到当前最优性能,并在新提出的MultiEditReward-Bench上超越专有评估器。此外,SpatialReward作为在线强化学习中的可靠信号,使OmniGen2在GEdit-Bench上提升+0.90,超过领先判别模型,且是GPT-4.1提升幅度(+0.45)的两倍。结果表明,空间推理对实现图像编辑的有效对齐至关重要。

原文摘要 · Abstract (English)

Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and fine-grained reward signals. Existing evaluators frequently struggle with a critical perception gap we term "Attention Collapse," where models neglect cross-image comparisons and fail to capture fine-grained details, resulting in inaccurate perception and miscalibrated scores. To address these limitations, we propose SpatialReward, a reward model that enforces precise verification via explicit spatial reasoning. By anchoring reasoning to predicted edit regions, SpatialReward grounds semantic judgments in pixel-level evidence, significantly enhancing evaluative accuracy. Trained on a curated 260k spatial-aware dataset, our model achieves state-of-the-art performance on MMRB2 and EditReward-Bench, and outperforms proprietary evaluators on our proposed MultiEditReward-Bench. Furthermore, SpatialReward serves as a robust signal in online RL, boosting OmniGen2 by +0.90 on GEdit-Bench--surpassing the leading discriminative model and doubling the gain of GPT-4.1 (+0.45). These results demonstrate that spatial reasoning is essential for unlocking effective alignment in image editing.

图像编辑强化学习空间推理奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。