arXiv:2409.10028cs.CVcs.AI2024-09

无需训练,通过调整注意力生成全新艺术风格。

AttnMod: Attention-Based New Art Styles

  • 利用注意力机制动态调节文本对图像的引导作用。
  • 可生成多种新颖风格,且不依赖提示词或模型重训。
  • 适合希望快速探索艺术风格的创作者与设计师。

我们提出 AttnMod,一种无需训练的技术,通过调节预训练扩散模型中的交叉注意力来生成全新的、无需提示词的艺术风格。该方法受人类艺术家重新诠释生成图像的启发,例如强调特定特征、分散色彩、扭曲轮廓或呈现未见元素。AttnMod 在去噪过程中改变文本提示对图像的注意力引导方式,实现针对性的风格调制。该方法无需修改提示词或重新训练模型,即可完成多样化的风格转换,显著扩展了文本到图像生成的表现力。

原文摘要 · Abstract (English)

We introduce AttnMod, a training-free technique that modulates cross-attention in pre-trained diffusion models to generate novel, unpromptable art styles. The method is inspired by how a human artist might reinterpret a generated image, for example by emphasizing certain features, dispersing color, twisting silhouettes, or materializing unseen elements. AttnMod simulates this intent by altering how the text prompt conditions the image through attention during denoising. These targeted modulations enable diverse stylistic transformations without changing the prompt or retraining the model, and they expand the expressive capacity of text-to-image generation.

艺术风格扩散模型注意力调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。