用CLIP+超网络让生成模型轻松适配新领域、文本操控图像。
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
- 通过超网络将CLIP空间融入预训练风格生成模型,实现动态适配。
- 无需额外训练数据即可实现文本引导的图像编辑与风格迁移。
- 新框架在多个任务上表现优于现有方法,尤其适合少样本场景。
生成对抗网络(GAN),特别是StyleGAN及其变体,在生成高度逼真图像方面表现出色。然而,将其适配到域适应、参考图像引导生成和文本引导操纵等多样化任务,尤其是在训练数据有限的情况下,仍具挑战性。为此,本文提出一种新框架,通过超网络将CLIP空间集成到预训练的StyleGAN中,实现对由参考图像或文本描述定义的新域的动态适应。此外,我们引入了基于CLIP的判别器,增强生成图像与目标域之间的对齐,确保高质量输出。该方法展现出前所未有的灵活性,可在无文本特定训练数据的情况下实现文本引导的图像操作,并支持无缝风格迁移。全面的定性和定量评估表明,本框架在性能和鲁棒性上均优于现有方法。
原文摘要 · Abstract (English)
Generative Adversarial Networks (GANs), particularly StyleGAN and its variants, have demonstrated remarkable capabilities in generating highly realistic images. Despite their success, adapting these models to diverse tasks such as domain adaptation, reference-guided synthesis, and text-guided manipulation with limited training data remains challenging. Towards this end, in this study, we present a novel framework that significantly extends the capabilities of a pre-trained StyleGAN by integrating CLIP space via hypernetworks. This integration allows dynamic adaptation of StyleGAN to new domains defined by reference images or textual descriptions. Additionally, we introduce a CLIP-guided discriminator that enhances the alignment between generated images and target domains, ensuring superior image quality. Our approach demonstrates unprecedented flexibility, enabling text-guided image manipulation without the need for text-specific training data and facilitating seamless style transfer. Comprehensive qualitative and quantitative evaluations confirm the robustness and superior performance of our framework compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。