让有瑕疵的图像也能生成高质量主体,自动去 artifacts
ArtiFade: Learning to Generate High-quality Subject from Blemished Images

- 用带瑕疵和无瑕疵图像对微调扩散模型,分离主体特征与干扰伪影
- 在分布内和分布外场景下均实现有效去伪影,生成质量显著提升
- 适合需要从低质输入生成高清图像的研究者或应用开发者
基于主体的文本到图像生成已能在少量图像下学习并捕捉主体特征。然而,现有方法通常依赖高质量图像训练,在输入图像存在伪影时难以生成合理结果,主要因当前技术难以区分主体特征与干扰伪影。本文提出 ArtiFade,通过微调预训练文本到图像模型,实现从带瑕疵数据集中生成高质量、无伪影图像。该方法利用包含无瑕疵图像及其对应带瑕疵版本的专用数据集进行微调,有效消除伪影,同时保留扩散模型原有的生成能力,从而提升主体驱动生成的整体性能。我们还设计了针对此任务的评估基准。通过大量定性和定量实验,证明 ArtiFade 在分布内与分布外场景下均具备良好的泛化性。
原文摘要 · Abstract (English)
Subject-driven text-to-image generation has witnessed remarkable advancements in its ability to learn and capture characteristics of a subject using only a limited number of images. However, existing methods commonly rely on high-quality images for training and may struggle to generate reasonable images when the input images are blemished by artifacts. This is primarily attributed to the inadequate capability of current techniques in distinguishing subject-related features from disruptive artifacts. In this paper, we introduce ArtiFade to tackle this issue and successfully generate high-quality artifact-free images from blemished datasets. Specifically, ArtiFade exploits fine-tuning of a pre-trained text-to-image model, aiming to remove artifacts. The elimination of artifacts is achieved by utilizing a specialized dataset that encompasses both unblemished images and their corresponding blemished counterparts during fine-tuning. ArtiFade also ensures the preservation of the original generative capabilities inherent within the diffusion model, thereby enhancing the overall performance of subject-driven methods in generating high-quality and artifact-free images. We further devise evaluation benchmarks tailored for this task. Through extensive qualitative and quantitative experiments, we demonstrate the generalizability of ArtiFade in effective artifact removal under both in-distribution and out-of-distribution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。