arXiv:2501.00816cs.CV2025-01中稿 · IEEE IEEE Transact…被引 9

无需训练即可提取多种艺术风格的草图,支持风格混合与精细控制。

MixSA: Training-free Reference-based Sketch Extraction via Mixture-of-Self-Attention

  • 用参考草图替换注意力机制中的键值,实现风格迁移。
  • 在解码器后期层调整轮廓,避免颜色混淆,提升纹理精度。
  • 可生成新风格草图,适合创意设计与艺术创作场景。

现有草图提取方法或需大量训练,或难以捕捉多样艺术风格,限制了实用性与灵活性。本文提出无需训练的混合自注意力(MixSA)方法,利用强大的扩散先验增强草图感知能力。其核心是将自注意力层中的键和值替换为参考草图的特征,使笔触元素无缝融入初始轮廓图,实现对纹理密度的精确控制,并支持风格插值以生成未见的新风格。通过在解码器晚期层对齐笔触风格与彩色图像的纹理和轮廓,有效缓解了因颜色平均化导致的失真问题。在多种感知指标下评估显示,MixSA在草图质量、灵活性和适用性方面均表现更优。该方法不仅克服了现有技术局限,还使用户能生成多样且高保真的草图,更准确体现丰富艺术表达。

原文摘要 · Abstract (English)

Current sketch extraction methods either require extensive training or fail to capture a wide range of artistic styles, limiting their practical applicability and versatility. We introduce Mixture-of-Self-Attention (MixSA), a training-free sketch extraction method that leverages strong diffusion priors for enhanced sketch perception. At its core, MixSA employs a mixture-of-self-attention technique, which manipulates self-attention layers by substituting the keys and values with those from reference sketches. This allows for the seamless integration of brushstroke elements into initial outline images, offering precise control over texture density and enabling interpolation between styles to create novel, unseen styles. By aligning brushstroke styles with the texture and contours of colored images, particularly in late decoder layers handling local textures, MixSA addresses the common issue of color averaging by adjusting initial outlines. Evaluated with various perceptual metrics, MixSA demonstrates superior performance in sketch quality, flexibility, and applicability. This approach not only overcomes the limitations of existing methods but also empowers users to generate diverse, high-fidelity sketches that more accurately reflect a wide range of artistic expressions.

草图生成风格迁移扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。