arXiv:2409.14720cs.CV2024-09被引 1

用文本和图像控制局部改衣,自监督训练不依赖真实数据集。

ControlEdit: A MultiModal Local Clothing Image Editing Method

  • 将服装编辑转为多模态引导的局部修复,实现精准修改。
  • 自监督学习解决真实数据难收集问题,保持编辑前后风格一致。
  • 自然过渡边界+非编辑区内容一致性,适合设计师快速试衣建模。

多模态服装图像编辑通过文本描述和视觉图像作为控制条件,实现对服装图像的精确调整与修改,显著提升设计师的工作效率并降低用户设计门槛。本文提出一种新方法 ControlEdit,将服装图像编辑任务转化为多模态引导的局部图像修复。针对真实图像数据集难以收集的问题,采用自监督学习方法进行训练。基于此,扩展特征提取网络通道数以保证编辑前后服装风格一致性,并设计逆潜在空间损失函数,实现对非编辑区域内容的软控制。此外,采用 Blended Latent Diffusion 作为采样方法,使编辑边界自然过渡,并强化非编辑区域内容的一致性。大量实验表明,ControlEdit 在定性和定量评估上均优于基线算法。

原文摘要 · Abstract (English)

Multimodal clothing image editing refers to the precise adjustment and modification of clothing images using data such as textual descriptions and visual images as control conditions, which effectively improves the work efficiency of designers and reduces the threshold for user design. In this paper, we propose a new image editing method ControlEdit, which transfers clothing image editing to multimodal-guided local inpainting of clothing images. We address the difficulty of collecting real image datasets by leveraging the self-supervised learning approach. Based on this learning approach, we extend the channels of the feature extraction network to ensure consistent clothing image style before and after editing, and we design an inverse latent loss function to achieve soft control over the content of non-edited areas. In addition, we adopt Blended Latent Diffusion as the sampling method to make the editing boundaries transition naturally and enforce consistency of non-edited area content. Extensive experiments demonstrate that ControlEdit surpasses baseline algorithms in both qualitative and quantitative evaluations.

图像编辑多模态自监督服装生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。