arXiv:2503.07307cs.CV2025-03被引 11

无需训练即可实现高质量风格迁移,兼顾内容保真与风格融合。

AttenST: A Training-Free Attention-Driven Style Transfer Framework with Pre-Trained Diffusion Models

  • 通过注意力机制融合内容与风格特征,不依赖微调或优化。
  • 在多个数据集上达到当前最优性能,显著提升风格保留效果。
  • 适合需要快速部署、低计算成本的图像风格迁移场景。

尽管扩散模型在风格迁移任务中取得显著进展,现有方法通常在推理时依赖微调或优化预训练模型,导致计算成本高,且难以平衡内容保持与风格融合。为此,我们提出AttenST,一种无需训练的注意力驱动风格迁移框架。具体而言,设计了一种风格引导的自注意力机制,保留内容图像的查询,但将键和值替换为风格图像的对应特征,实现有效风格特征融合。为缓解反演过程中的风格信息损失,引入风格保持型反演策略,通过多轮重采样提升反演精度。此外,提出内容感知的自适应实例归一化,将内容统计融入归一化过程,优化风格融合并减少内容退化。进一步引入双特征交叉注意力机制,融合内容与风格特征,确保结构保真与风格表达的和谐统一。大量实验表明,AttenST优于现有方法,在风格迁移数据集上达到最先进性能。

原文摘要 · Abstract (English)

While diffusion models have achieved remarkable progress in style transfer tasks, existing methods typically rely on fine-tuning or optimizing pre-trained models during inference, leading to high computational costs and challenges in balancing content preservation with style integration. To address these limitations, we introduce AttenST, a training-free attention-driven style transfer framework. Specifically, we propose a style-guided self-attention mechanism that conditions self-attention on the reference style by retaining the query of the content image while substituting its key and value with those from the style image, enabling effective style feature integration. To mitigate style information loss during inversion, we introduce a style-preserving inversion strategy that refines inversion accuracy through multiple resampling steps. Additionally, we propose a content-aware adaptive instance normalization, which integrates content statistics into the normalization process to optimize style fusion while mitigating the content degradation. Furthermore, we introduce a dual-feature cross-attention mechanism to fuse content and style features, ensuring a harmonious synthesis of structural fidelity and stylistic expression. Extensive experiments demonstrate that AttenST outperforms existing methods, achieving state-of-the-art performance in style transfer dataset.

风格迁移扩散模型注意力机制无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。