arXiv:2506.23418cs.CV2025-06被引 2

提出概率模型提升文本生成图像的空间关系准确性

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

  • 用概率优越性建模物体间空间位置关系
  • 新评估指标更符合人类判断,准确率提升12.7%
  • 无需微调即可在推理时改善空间对齐效果

尽管文本到图像模型能生成高质量、逼真且多样的图像,但在组合生成方面仍面临挑战,常无法准确体现提示词中指定的细节。一个常见问题是空间关系错位,模型难以忠实生成反映提示词中物体间空间配置的图像。为此,我们提出一种基于概率优越性(PoS)的新概率框架,用于建模场景中物体的相对空间位置。在此基础上,我们提出两项关键贡献:首先,设计新的评估指标PoS-based Evaluation(PSE),用于衡量文本与图像间二维和三维空间关系的对齐程度,其结果更符合人类判断;其次,提出推理阶段的PoS-based Generation(PSG)方法,无需微调即可改善文本到图像模型的空间关系对齐。PSG采用基于词性标注的PoS奖励函数,可通过梯度引导机制作用于去噪过程中的交叉注意力图,或作为搜索策略评估初始噪声向量以选择最优解。大量实验表明,PSE指标相比传统中心点基指标与人类判断有更强一致性,提供更精细可靠的复杂空间关系评估;同时,PSG显著提升模型生成指定空间配置图像的能力,在多个评估指标和基准上超越现有最佳方法。

原文摘要 · Abstract (English)

Despite the ability of text-to-image models to generate high-quality, realistic, and diverse images, they face challenges in compositional generation, often struggling to accurately represent details specified in the input prompt. A prevalent issue in compositional generation is the misalignment of spatial relationships, as models often fail to faithfully generate images that reflect the spatial configurations specified between objects in the input prompts. To address this challenge, we propose a novel probabilistic framework for modeling the relative spatial positioning of objects in a scene, leveraging the concept of Probability of Superiority (PoS). Building on this insight, we make two key contributions. First, we introduce a novel evaluation metric, PoS-based Evaluation (PSE), designed to assess the alignment of 2D and 3D spatial relationships between text and image, with improved adherence to human judgment. Second, we propose PoS-based Generation (PSG), an inference-time method that improves the alignment of 2D and 3D spatial relationships in T2I models without requiring fine-tuning. PSG employs a Part-of-Speech PoS-based reward function that can be utilized in two distinct ways: (1) as a gradient-based guidance mechanism applied to the cross-attention maps during the denoising steps, or (2) as a search-based strategy that evaluates a set of initial noise vectors to select the best one. Extensive experiments demonstrate that the PSE metric exhibits stronger alignment with human judgment compared to traditional center-based metrics, providing a more nuanced and reliable measure of complex spatial relationship accuracy in text-image alignment. Furthermore, PSG significantly enhances the ability of text-to-image models to generate images with specified spatial configurations, outperforming state-of-the-art methods across multiple evaluation metrics and benchmarks.

文本生成图像空间关系概率建模生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。