用单图实现精准图像定制,分离主体与干扰信息
DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation
- 通过分离视觉提示中的主体与无关成分,生成专属特征
- 在无需微调情况下实现高保真度身份一致性和编辑能力
- 适合需要快速个性化图像生成的场景
在图像生成领域,基于视觉提示结合文本指令生成定制化图像成为重要方向。然而,现有方法(无论是否微调)难以准确提取视觉提示中的主体关键属性,导致无关属性渗入生成过程,降低个性化质量,影响可编辑性与身份一致性。本文提出DisEnvisioner,一种无需微调、仅需单张图像即可有效提取并增强主体关键特征的新方法。该方法将主体特征与其他无关成分分离为独立视觉标记,实现更精准的定制。为进一步提升身份一致性,对解耦特征进行精细化增强,构建更细粒度表示。实验表明,DisEnvisioner在指令响应(可编辑性)、身份一致性、推理速度和整体图像质量方面均优于现有方法,充分验证其有效性与高效性。
原文摘要 · Abstract (English)
In the realm of image generation, creating customized images from visual prompt with additional textual instruction emerges as a promising endeavor. However, existing methods, both tuning-based and tuning-free, struggle with interpreting the subject-essential attributes from the visual prompt. This leads to subject-irrelevant attributes infiltrating the generation process, ultimately compromising the personalization quality in both editability and ID preservation. In this paper, we present DisEnvisioner, a novel approach for effectively extracting and enriching the subject-essential features while filtering out -irrelevant information, enabling exceptional customization performance, in a tuning-free manner and using only a single image. Specifically, the feature of the subject and other irrelevant components are effectively separated into distinctive visual tokens, enabling a much more accurate customization. Aiming to further improving the ID consistency, we enrich the disentangled features, sculpting them into more granular representations. Experiments demonstrate the superiority of our approach over existing methods in instruction response (editability), ID consistency, inference speed, and the overall image quality, highlighting the effectiveness and efficiency of DisEnvisioner. Project page: https://disenvisioner.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。