arXiv:2510.01347cs.CV2025-10

用一张参考图提取风格,精准控制生成图像风格。

Image Generation Based on Image Style Extraction

  • 从单张参考图提取细粒度风格特征,注入生成模型。
  • 在不改动模型结构下实现文本与风格的精准对齐。
  • 适合需要精细风格控制的研究者和设计师。

基于文本的图像生成在实际应用中面临挑战:细粒度风格难以用自然语言精确描述和控制,而风格参考图像的指导信息又难以与传统文本引导生成对齐。本研究旨在最大化预训练生成模型的生成能力,通过从单张风格参考图中提取细粒度风格表征,并将其无修改地注入生成主体,实现细粒度可控的风格化图像生成。提出一种三阶段训练的风格提取生成方法,采用风格编码器与风格投影层,将风格表征与文本表征对齐,实现基于文本提示的风格引导生成。同时构建了包含图像、风格标签与文本描述三元组的Style30k-captions数据集,用于训练风格编码器与投影层。

原文摘要 · Abstract (English)

Image generation based on text-to-image generation models is a task with practical application scenarios that fine-grained styles cannot be precisely described and controlled in natural language, while the guidance information of stylized reference images is difficult to be directly aligned with the textual conditions of traditional textual guidance generation. This study focuses on how to maximize the generative capability of the pretrained generative model, by obtaining fine-grained stylistic representations from a single given stylistic reference image, and injecting the stylistic representations into the generative body without changing the structural framework of the downstream generative model, so as to achieve fine-grained controlled stylized image generation. In this study, we propose a three-stage training style extraction-based image generation method, which uses a style encoder and a style projection layer to align the style representations with the textual representations to realize fine-grained textual cue-based style guide generation. In addition, this study constructs the Style30k-captions dataset, whose samples contain a triad of images, style labels, and text descriptions, to train the style encoder and style projection layer in this experiment.

风格迁移图像生成文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。