arXiv:2511.06365cs.CV2025-11

通过打乱注意力值特征,实现零样本风格迁移且避免内容泄露。

V-Shuffle: Zero-Shot Style Transfer via Value Shuffle

  • 在扩散模型自注意力层打乱值特征,抑制风格图像语义干扰。
  • 多风格图输入时显著提升风格保真度,单图也优于现有方法。
  • 适合需要高保真、无内容污染风格迁移的场景。

基于注意力注入的风格迁移近年来取得显著进展,但现有方法常出现内容泄露问题,即风格图像中不期望的语义内容错误地出现在生成结果中。本文提出V-Shuffle,一种零样本风格迁移方法,利用同一风格域内的多张风格图像,有效平衡内容保留与风格保真之间的权衡。V-Shuffle通过在扩散模型的自注意力层内打乱值特征,隐式破坏风格图像的语义内容,从而保留低层次风格表征。我们进一步引入混合风格正则化,结合高层次风格纹理以增强风格保真度。实验表明,当使用多张风格图像时,V-Shuffle表现优异;而仅使用单张风格图像时,其性能亦超越先前最先进方法。

原文摘要 · Abstract (English)

Attention injection-based style transfer has achieved remarkable progress in recent years. However, existing methods often suffer from content leakage, where the undesired semantic content of the style image mistakenly appears in the stylized output. In this paper, we propose V-Shuffle, a zero-shot style transfer method that leverages multiple style images from the same style domain to effectively navigate the trade-off between content preservation and style fidelity. V-Shuffle implicitly disrupts the semantic content of the style images by shuffling the value features within the self-attention layers of the diffusion model, thereby preserving low-level style representations. We further introduce a Hybrid Style Regularization that complements these low-level representations with high-level style textures to enhance style fidelity. Empirical results demonstrate that V-Shuffle achieves excellent performance when utilizing multiple style images. Moreover, when applied to a single style image, V-Shuffle outperforms previous state-of-the-art methods.

风格迁移扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。