无需训练即可融合多风格图像,实现可控的视觉风格混合。
Training-Free Multi-Style Fusion Through Reference-Based Adaptive Modulation
- 通过语义标记分解模块注入多个风格参考图。
- 每步去噪时动态调整各风格注意力权重,实现平衡融合。
- 无需微调,支持任意数量风格混合,适合创意设计场景。
我们提出自适应多风格融合(AMSF),一种基于参考图像的无训练框架,可在扩散模型中可控地融合多个参考风格。现有方法通常仅支持单张风格图,难以实现混合美学或扩展至多风格;且缺乏平衡多种风格影响的合理机制。AMSF通过语义标记分解模块编码所有风格图像与文本提示,并自适应注入冻结扩散模型的每个交叉注意力层。随后,相似度感知重加权模块在每步去噪时重新校准各风格成分的注意力分配,实现无需微调或外部适配器的平衡、可调控融合。定性与定量评估均表明,AMSF生成结果持续优于当前最优方法,其融合设计可无缝扩展至两个及以上风格。该能力为扩散模型中的表达性多风格生成提供了实用路径。
原文摘要 · Abstract (English)
We propose Adaptive Multi-Style Fusion (AMSF), a reference-based training-free framework that enables controllable fusion of multiple reference styles in diffusion models. Most of the existing reference-based methods are limited by (a) acceptance of only one style image, thus prohibiting hybrid aesthetics and scalability to more styles, and (b) lack of a principled mechanism to balance several stylistic influences. AMSF mitigates these challenges by encoding all style images and textual hints with a semantic token decomposition module that is adaptively injected into every cross-attention layer of an frozen diffusion model. A similarity-aware re-weighting module then recalibrates, at each denoising step, the attention allocated to every style component, yielding balanced and user-controllable blends without any fine-tuning or external adapters. Both qualitative and quantitative evaluations show that AMSF produces multi-style fusion results that consistently outperform the state-of-the-art approaches, while its fusion design scales seamlessly to two or more styles. These capabilities position AMSF as a practical step toward expressive multi-style generation in diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。