不训练即可精准对齐图文位置,解决生成图像错位问题
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
- 用最优传输理论动态调整注意力图,实现空间位置精准控制
- 早期去噪阶段引入空间感知,显著减少物体错位现象
- 无需训练,适合快速适配新任务,提升生成图像空间一致性
基于扩散的文本到图像(T2I)模型在无需训练的情况下已实现高质量图像生成,具备低成本适应与跨任务泛化能力。然而,现有方法虽持续应对‘物体缺失’和‘属性错配’等问题,却仍难以解决‘物体错位’这一关键挑战:生成图像中的物体位置无法准确对应文本提示。这源于文本形式难以提供显式空间引导。为此,我们提出STORM(Spatial Transport Optimization by Repositioning Attention Map),一种新的无需训练的空间一致T2I生成方法。STORM采用基于最优传输理论的时空优化(STO),通过空间传输(ST)代价函数动态调整注意力图,增强空间理解。分析表明,空间感知在早期去噪阶段最有效,后期则用于细节优化。大量实验显示,STORM在缓解物体错位的同时,也改善了缺失与属性错配问题,成为当前T2I生成中空间对齐的新基准。
原文摘要 · Abstract (English)
Diffusion-based text-to-image (T2I) models have recently excelled in high-quality image generation, particularly in a training-free manner, enabling cost-effective adaptability and generalization across diverse tasks. However, while the existing methods have been continuously focusing on several challenges, such as "missing objects" and "mismatched attributes," another critical issue of "mislocated objects" remains where generated spatial positions fail to align with text prompts. Surprisingly, ensuring such seemingly basic functionality remains challenging in popular T2I models due to the inherent difficulty of imposing explicit spatial guidance via text forms. To address this, we propose STORM (Spatial Transport Optimization by Repositioning Attention Map), a novel training-free approach for spatially coherent T2I synthesis. STORM employs Spatial Transport Optimization (STO), rooted in optimal transport theory, to dynamically adjust object attention maps for precise spatial adherence, supported by a Spatial Transport (ST) Cost function that enhances spatial understanding. Our analysis shows that integrating spatial awareness is most effective in the early denoising stages, while later phases refine details. Extensive experiments demonstrate that STORM surpasses existing methods, effectively mitigating mislocated objects while improving missing and mismatched attributes, setting a new benchmark for spatial alignment in T2I synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。