无需训练,任意参考图也能精准个性化风格迁移
StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References
- 基于隐空间聚类实现无监督区域分割,不依赖额外标注
- 多参考图输入下内容结构保留率提升18%,局部风格更精准
- 适合追求自由风格定制的设计师和创意工作者
尽管基于扩散模型的图像风格迁移取得进展,现有方法仍受限于:1)语义鸿沟——风格参考缺乏恰当内容语义,导致不可控的风格化;2)依赖额外约束(如语义掩码),限制适用性;3)特征关联僵化,缺乏自适应的全局-局部对齐,难以兼顾细粒度风格与整体内容保真。这些局限尤其体现在无法灵活利用风格输入,严重制约了风格迁移在个性化、准确性与适应性方面的表现。为此,我们提出StyleGallery,一个无需训练且具备语义感知能力的框架,支持任意参考图像作为输入,实现高效的个性化定制。其包含三个核心阶段:语义区域分割(在扩散隐空间中进行自适应聚类以划分区域,无需额外输入);聚类区域匹配(对提取特征进行块级过滤以实现精确对齐);风格迁移优化(通过能量函数引导的扩散采样与区域风格损失优化风格化)。在我们提出的基准测试上,实验表明StyleGallery在内容结构保留、区域风格化、可解释性及个性化定制方面均优于现有最先进方法,尤其是在使用多个风格参考时表现突出。
原文摘要 · Abstract (English)
Despite the advancements in diffusion-based image style transfer, existing methods are commonly limited by 1) semantic gap: the style reference could miss proper content semantics, causing uncontrollable stylization; 2) reliance on extra constraints (e.g., semantic masks) restricting applicability; 3) rigid feature associations lacking adaptive global-local alignment, failing to balance fine-grained stylization and global content preservation. These limitations, particularly the inability to flexibly leverage style inputs, fundamentally restrict style transfer in terms of personalization, accuracy, and adaptability. To address these, we propose StyleGallery, a training-free and semantic-aware framework that supports arbitrary reference images as input and enables effective personalized customization. It comprises three core stages: semantic region segmentation (adaptive clustering on latent diffusion features to divide regions without extra inputs); clustered region matching (block filtering on extracted features for precise alignment); and style transfer optimization (energy function-guided diffusion sampling with regional style loss to optimize stylization). Experiments on our introduced benchmark demonstrate that StyleGallery outperforms state-of-the-art methods in content structure preservation, regional stylization, interpretability, and personalized customization, particularly when leveraging multiple style references.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。