让生成图像精准保留参考图内容,同时自由变换风格。
Content-Aware Preserving Image Generation
- 用频域特征提取+内容融合模块,生成可控内容的图像
- 在多个数据集上实现高保真内容还原,风格多样且一致
- 适合需要精确控制图像内容的创意设计与视觉生成
生成模型在图像生成任务中取得了显著进展,但精确控制生成图像的内容仍具挑战性,源于其固有的训练目标。本文提出一种新型图像生成框架,旨在显式融入期望内容。该框架采用先进编码技术,集成内容融合与频率编码子网络。频率编码模块通过聚焦特定频段成分,捕获参考图像的特征与结构;内容融合模块则生成包含目标内容特征的内容引导向量。在生成过程中,真实图像的内容引导向量与投影噪声向量融合,确保生成图像既保持指导图像的内容一致性,又呈现多样化风格变化。为验证框架在内容保持方面的有效性,我们在Flickr-Faces-High Quality、Animal Faces High Quality和Large-scale Scene Understanding等广泛使用的基准数据集上进行了大量实验。
原文摘要 · Abstract (English)
Remarkable progress has been achieved in image generation with the introduction of generative models. However, precisely controlling the content in generated images remains a challenging task due to their fundamental training objective. This paper addresses this challenge by proposing a novel image generation framework explicitly designed to incorporate desired content in output images. The framework utilizes advanced encoding techniques, integrating subnetworks called content fusion and frequency encoding modules. The frequency encoding module first captures features and structures of reference images by exclusively focusing on selected frequency components. Subsequently, the content fusion module generates a content-guiding vector that encapsulates desired content features. During the image generation process, content-guiding vectors from real images are fused with projected noise vectors. This ensures the production of generated images that not only maintain consistent content from guiding images but also exhibit diverse stylistic variations. To validate the effectiveness of the proposed framework in preserving content attributes, extensive experiments are conducted on widely used benchmark datasets, including Flickr-Faces-High Quality, Animal Faces High Quality, and Large-scale Scene Understanding datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。