arXiv:2509.15962cs.AI2025-09被引 1

用结构化信息提升文本生成图像的方位准确性

Structured Information for Improving Spatial Relationships in Text-to-Image Generation

  • 用细调语言模型自动生成提示词中的空间关系元组
  • 空间准确率显著提升,图像质量无损失
  • 适合需要精准布局的图像生成场景

文本到图像(T2I)生成已迅速发展,但准确捕捉自然语言提示中的空间关系仍是主要挑战。现有方法通过提示优化、空间对齐生成和语义精炼来应对。本文提出一种轻量级方案,利用微调语言模型将提示词自动转换为基于元组的结构化信息,并无缝集成到T2I流程中。实验表明,该方法在不降低图像整体质量(以Inception Score衡量)的前提下,显著提升了空间准确性。此外,自动生成的元组质量与人工构建相当。这种结构化信息为增强T2I生成中的空间关系提供了实用且可移植的解决方案,有效缓解了当前大规模生成系统的关键缺陷。

原文摘要 · Abstract (English)

Text-to-image (T2I) generation has advanced rapidly, yet faithfully capturing spatial relationships described in natural language prompts remains a major challenge. Prior efforts have addressed this issue through prompt optimization, spatially grounded generation, and semantic refinement. This work introduces a lightweight approach that augments prompts with tuple-based structured information, using a fine-tuned language model for automatic conversion and seamless integration into T2I pipelines. Experimental results demonstrate substantial improvements in spatial accuracy, without compromising overall image quality as measured by Inception Score. Furthermore, the automatically generated tuples exhibit quality comparable to human-crafted tuples. This structured information provides a practical and portable solution to enhance spatial relationships in T2I generation, addressing a key limitation of current large-scale generative systems.

文本生成图像空间关系结构化提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。