arXiv:2412.03812cs.CV2024-12

Pinco通过位置感知适配器,让图像修复更贴合文字描述且不扭曲主体形状。

Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned Inpainting

  • 在自注意力层融合主体特征,缓解文本与图像冲突
  • 分离语义与空间特征提取,更好保留主体轮廓
  • 共享位置嵌入锚点提升定位精度,适合需要高保真修复的场景

前景条件图像修复旨在利用给定的前景主体和文本描述无缝填充图像背景。现有基于文本到图像生成的方法在此任务中存在主体形状扩展、失真或与文本描述对齐能力下降的问题,导致视觉元素与文本描述不一致。为此,我们提出Pinco,一种即插即用的前景条件修复适配器,在保持主体形状的同时生成高质量背景并实现良好文本对齐。首先,设计自一致性适配器,将前景特征融入布局相关的自注意力层,使模型在处理整体图像布局时能有效考虑前景特征,缓解文本与主体特征的冲突。其次,设计解耦图像特征提取方法,采用不同架构分别提取语义与空间特征,显著提升主体特征提取质量,保障主体形状的高保真性。第三,引入共享位置嵌入锚点,精确利用提取特征并聚焦于主体区域,大幅提升模型对主体特征的理解能力,同时提升训练效率。大量实验表明,该方法在前景条件修复任务中表现优异且高效。

原文摘要 · Abstract (English)

Foreground-conditioned inpainting aims to seamlessly fill the background region of an image by utilizing the provided foreground subject and a text description. While existing T2I-based image inpainting methods can be applied to this task, they suffer from issues of subject shape expansion, distortion, or impaired ability to align with the text description, resulting in inconsistencies between the visual elements and the text description. To address these challenges, we propose Pinco, a plug-and-play foreground-conditioned inpainting adapter that generates high-quality backgrounds with good text alignment while effectively preserving the shape of the foreground subject. Firstly, we design a Self-Consistent Adapter that integrates the foreground subject features into the layout-related self-attention layer, which helps to alleviate conflicts between the text and subject features by ensuring that the model can effectively consider the foreground subject's characteristics while processing the overall image layout. Secondly, we design a Decoupled Image Feature Extraction method that employs distinct architectures to extract semantic and spatial features separately, significantly improving subject feature extraction and ensuring high-quality preservation of the subject's shape. Thirdly, to ensure precise utilization of the extracted features and to focus attention on the subject region, we introduce a Shared Positional Embedding Anchor, greatly improving the model's understanding of subject features and boosting training efficiency. Extensive experiments demonstrate that our method achieves superior performance and efficiency in foreground-conditioned inpainting.

图像修复扩散模型文本对齐主体保形

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。