arXiv:2606.10892cs.CVcs.AI2026-06

用定制化概念嵌入提升图像外扩的文本-实例对齐,减少合成伪影。

Improving Text-Instance Alignment Of Foreground Conditioned Out-Painting Via Customized Concept Embedding

论文配图:Improving Text-Instance Alignment Of Foreground Conditioned Out-Painting Via Customized Concept Embedding
图 1 · 摘自论文原文
  • 设计定制化概念嵌入模块,精准对齐文本与具体物体视觉特征。
  • 在多个数据集上显著降低伪影率,提升背景合成质量。
  • 可插拔集成,适合需要高质量商品图生成的研究与应用。

为展示商品,商家常需耗费大量成本制作高质量展示图。前景条件外扩(FCO)通过调整文本提示,低成本生成目标背景,满足需求。然而现有文本驱动的FCO方法存在严重缺陷,主要表现为合成背景中出现与前景实例语义相同的伪影区域,削弱主体突出性并降低图像质量。我们归因于实例与文本生成的概念嵌入之间对齐不足。为此,提出定制化概念嵌入扩散框架(CCE-Diffusion),其核心是CCE模块,用于定制化概念嵌入,弥合通用名词语义与特定视觉实例之间的差距。实例感知损失指导模块优化,语义保持提示模板防止定制嵌入扭曲提示中其他词汇。定性与定量评估均表明,CCE-Diffusion显著减少输出中的伪影。作为即插即用组件,CCE模块可集成至多种FCO方法中,提升其性能。

原文摘要 · Abstract (English)

To showcase products, merchants often incur substantial costs creating high-quality display images. Foreground Conditioned Outpainting (FCO) meets this demand, allowing users to create desired backgrounds for foreground instances at a low cost by adjusting the text prompt. However, existing text-driven FCO methods exhibit critical flaws in their outputs, most notably the presence of artifacts, which refer to regions in the synthesized background that share the same semantics as the foreground instance. Such artifacts diminish the object's prominence and degrade image quality. We attribute the issue to the misalignment between the given instance and text-derived concept embeddings. To address this, we propose the Customized Concept Embedding Diffusion (CCE-Diffusion) framework. Its core is a CCE-Module to customize concept embeddings, bridging the gap between generic noun semantics and a specific visual instance. An Instance-Aware Loss guides the module's optimization, while a Semantic-Preserving Prompt Template prevents customized embeddings from distorting other words in the prompt. Both qualitative and quantitative evaluations demonstrate that CCE-Diffusion significantly reduces artifacts in the outputs. As a plug-and-play component, the CCE-Module can integrate with various FCO methods, enhancing their performance.

图像生成文本对齐扩散模型伪影消除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。