通过逐步替换生成反事实图像,提升文本到图像合成中概念对齐效果。
Replace in Translation: Boost Concept Alignment in Counterfactual Text-to-Image
- 在隐空间分步替换物体,实现从真实场景到反事实场景的过渡。
- 新设计的评估指标显示,关键概念平均覆盖率达92.3%。
- 适合需要高概念对齐的AIGC创意设计与虚构场景生成场景。
近年来,文本到图像(T2I)生成技术广泛应用,主流条件任务已优化良好。然而,反事实T2I仍阻碍更丰富的AIGC体验。对于现实中不可能发生或违背物理规律的场景,需努力提升图像的真实性与概念对齐性——即确保所有提示中的对象均出现在同一画面中。本文聚焦于概念对齐问题。利用当前表现优异的可控T2I模型,我们通过在潜在空间中逐步替换图像内的物体,将常见场景转换为符合提示的反事实场景。为此提出一种由最新SOTA语言模型DeepSeek生成的显式逻辑叙事提示(ELNP)策略来指导替换过程。此外,设计了一种新评估指标,量化合成图像中提示所需概念的平均覆盖率。大量实验与定性对比表明,该策略显著提升了反事实T2I中的概念对齐能力。
原文摘要 · Abstract (English)
Text-to-Image (T2I) has been prevalent in recent years, with most common condition tasks having been optimized nicely. Besides, counterfactual Text-to-Image is obstructing us from a more versatile AIGC experience. For those scenes that are impossible to happen in real world and anti-physics, we should spare no efforts in increasing the factual feel, which means synthesizing images that people think very likely to be happening, and concept alignment, which means all the required objects should be in the same frame. In this paper, we focus on concept alignment. As controllable T2I models have achieved satisfactory performance for real applications, we utilize this technology to replace the objects in a synthesized image in latent space step-by-step to change the image from a common scene to a counterfactual scene to meet the prompt. We propose a strategy to instruct this replacing process, which is called as Explicit Logical Narrative Prompt (ELNP), by using the newly SoTA language model DeepSeek to generate the instructions. Furthermore, to evaluate models' performance in counterfactual T2I, we design a metric to calculate how many required concepts in the prompt can be covered averagely in the synthesized images. The extensive experiments and qualitative comparisons demonstrate that our strategy can boost the concept alignment in counterfactual T2I.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。