提出可验证的空间奖励模型,让图文生成更准确定位物体位置。
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
- 从提示词中提取实体、属性和空间信息,分步评估图像空间布局
- 在Stable Diffusion和FLUX上提升空间一致性,接近人类判断
- 适合关注精准定位与复杂关系生成的研究者
基于强化学习的文本到图像生成近年受益于评估语义对齐与视觉质量的奖励模型。然而,现有方法对细粒度空间关系关注不足,常导致图像整体看似合理但物体位置存在偏差。本文提出SpatialReward,一种可验证的空间奖励模型,专门评估生成图像中的空间布局。该模型采用多阶段流程:提示分解器从自由文本中提取实体、属性与空间元数据;专家检测器提供物体位置与属性的精准视觉定位;视觉语言模型基于定位结果进行链式推理,评估规则方法难以处理的复杂空间关系。为全面评估生成图像中的空间关系,我们构建SpatRelBench基准,涵盖物体属性、朝向、物体间关系及文本渲染位置。在Stable Diffusion和FLUX上的实验表明,将SpatialReward引入强化学习训练后,空间一致性和整体生成质量显著提升,结果更贴近人类判断。这表明可验证的奖励模型在实现更精准可控的图文生成优化方面具有巨大潜力。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to fine-grained spatial relationships, often producing images that appear plausible overall yet contain inaccuracies in object positioning. In this work, we present \textbf{SpatialReward}, a verifiable reward model explicitly designed to evaluate spatial layouts in generated images. SpatialReward adopts a multi-stage pipeline: a \emph{Prompt Decomposer} extracts entities, attributes, and spatial metadata from free-form prompts; expert detectors provide accurate visual grounding of object positions and attributes; and a vision-language model applies chain-of-thought reasoning over grounded observations to assess complex spatial relations that are challenging for rule-based methods. To more comprehensively evaluate spatial relationships in generated images, we introduce \textbf{SpatRelBench}, a benchmark covering object attributes, orientation, inter-object relations, and rendered text placement. Experiments on Stable Diffusion and FLUX show that incorporating SpatialReward into RL training consistently improves spatial consistency and overall generation quality, with results aligned more closely to human judgments. These findings indicate that verifiable reward models hold considerable potential for enabling more accurate and controllable optimization in text-to-image generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。