提升文生图模型的空间精度,不增加生成时间。
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
- 基于低秩适配的灵活微调框架,增强图像空间一致性。
- 在空间一致性评测中优于现有基线方法CoMPaSS。
- 提出新评估指标,可发现并利用模型空间偏见。
扩散模型已革新文生图合成,生成高质量、逼真图像。然而,它们仍难以准确呈现文本提示中的空间关系。现有方法多依赖外部网络条件或预设布局,导致计算开销大、灵活性差。本文构建了一个精心筛选的空间显式提示数据集,从LAION-400M中提取并合成,确保文本描述与空间布局精准对齐。同时提出ESPLoRA,一种基于低秩适配的灵活微调框架,可在不增加生成时间的前提下提升生成模型的空间一致性,且不影响输出质量。此外,我们设计了基于几何约束的改进评估指标,能捕捉如“在...前面”或“在...后面”等3D空间关系。这些指标揭示了文生图模型中的空间偏见,我们进一步通过TORE算法策略性利用该偏见,显著提升生成图像的空间一致性。实验表明,本方法在空间一致性基准上优于当前基线框架CoMPaSS。
原文摘要 · Abstract (English)
Diffusion models have revolutionized text-to-image (T2I) synthesis, producing high-quality, photorealistic images. However, they still struggle to properly render the spatial relationships described in text prompts. To address the lack of spatial information in T2I generations, existing methods typically use external network conditioning and predefined layouts, resulting in higher computational costs and reduced flexibility. Our approach builds upon a curated dataset of spatially explicit prompts, meticulously extracted and synthesized from LAION-400M to ensure precise alignment between textual descriptions and spatial layouts. Alongside this dataset, we present ESPLoRA, a flexible fine-tuning framework based on Low-Rank Adaptation, specifically designed to enhance spatial consistency in generative models without increasing generation time or compromising the quality of the outputs. In addition to ESPLoRA, we propose refined evaluation metrics grounded in geometric constraints, capturing 3D spatial relations such as "in front of" or "behind". These metrics also expose spatial biases in T2I models which, even when not fully mitigated, can be strategically exploited by our TORE algorithm to further improve the spatial consistency of generated images. Our method outperforms CoMPaSS, the current baseline framework, on spatial consistency benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。