用多幅画作学习艺术家整体风格,生成更真实的作品。
Through Van Gogh's Eyes: Global Style Transfer with Diffusion Model

- 用多幅作品提取艺术家全局风格,非单一作品模仿
- 生成作品风格还原度高,内容结构保留好,多样性更强
- 适合想生成真实艺术风格作品的研究者和创作者
艺术图像合成旨在再现目标艺术家的表达性视觉特征,但现有方法常无法捕捉艺术家的整体风格。传统风格迁移方法以一对一方式将单幅参考作品风格迁移至内容图像,虽适用于作品级风格化,却难以体现艺术家整体风格分布。基于艺术家名称的文本到图像扩散模型(如“~梵高风格”)虽具灵活性,但易受文本偏见影响,仅复现少数经典作品的模式。为此,我们提出全局风格迁移(GST),以多对一方式聚合目标艺术家的多幅作品,将其共享的全局风格迁移到单个内容图像。针对GST,我们提出全局风格引导(GSG),在固定提示下于扩散模型的中间特征空间(h-space)中学习残差全局风格偏移,仅通过视觉统计学习艺术家级风格语义,消除文本依赖偏差。进一步提出内容对齐引导(CAG),一种无需训练的感知引导机制,在保持内容语义结构的同时允许特定艺术家的几何变形。在WikiArt上的实验表明,GST在风格保真度、内容保留与输出多样性方面均优于现有风格迁移及基于扩散的艺术合成方法。
原文摘要 · Abstract (English)
Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for artwork-level stylization but limited in representing the broader stylistic distribution of an artist. Text-to-image diffusion models conditioned on artist names, such as '~ in Van Gogh style', offer greater flexibility, but they often suffer from text-induced bias and reproduce patterns from only a few iconic works. To address these limitations, we introduce Global Style Transfer (GST), an artistic image synthesis paradigm, in a Many-to-One manner, that aggregates multiple artworks from a target artist and transfers their shared global style to a single content image. For GST, we propose Global Style Guidance (GSG), which learns a residual global style offset in the intermediate feature space, or h-space, of a diffusion model under a fixed prompt. By learning artist-level style semantics purely from visual statistics, GSG mitigates text-dependent artistic bias. We further propose Content Alignment Guidance (CAG), a training-free perceptual guidance mechanism that preserves the semantic structure of the content image while allowing artist-specific geometric deformation. Experiments on WikiArt demonstrate that GST achieves superior stylistic fidelity, content preservation, and output diversity compared to existing style transfer and diffusion-based artistic synthesis methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。