arXiv:2602.24233cs.CV2026-02中稿 · CVPR被引 1

用奖励模型提升图像生成的空间关系理解能力。

Enhancing Spatial Understanding in Image Generation via Reward Modeling

  • 构建8万+偏好对数据集,训练空间关系评分模型。
  • 在多个基准上显著提升图像生成的空间准确性。
  • 适合需要精准空间布局的文本生成图像任务。

文本到图像生成近年在视觉保真度和创造力方面取得显著进展,但也对提示复杂度提出了更高要求,尤其在表达复杂空间关系时,常需多次采样才能获得满意结果。为解决此问题,我们提出一种新方法,增强现有图像生成模型的空间理解能力。首先构建包含8万余个偏好对的SpatialReward-Dataset。基于该数据集,我们设计了SpatialScore奖励模型,用于评估文本到图像生成中空间关系的准确性,在空间评价任务上表现甚至超越领先专有模型。进一步实验表明,该奖励模型可有效支持复杂空间生成的在线强化学习。在多个基准上的大量实验显示,我们的专用奖励模型在图像生成的空间理解上带来了显著且一致的提升。

原文摘要 · Abstract (English)

Recent progress in text-to-image generation has greatly advanced visual fidelity and creativity, but it has also imposed higher demands on prompt complexity-particularly in encoding intricate spatial relationships. In such cases, achieving satisfactory results often requires multiple sampling attempts. To address this challenge, we introduce a novel method that strengthens the spatial understanding of current image generation models. We first construct the SpatialReward-Dataset with over 80k preference pairs. Building on this dataset, we build SpatialScore, a reward model designed to evaluate the accuracy of spatial relationships in text-to-image generation, achieving performance that even surpasses leading proprietary models on spatial evaluation. We further demonstrate that this reward model effectively enables online reinforcement learning for the complex spatial generation. Extensive experiments across multiple benchmarks show that our specialized reward model yields significant and consistent gains in spatial understanding for image generation.

图像生成空间理解奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。