arXiv:2604.12281cs.CVcs.AI2026-04

无需训练即可实现多风格迁移,保持结构清晰无伪影。

MAST: Mask-Guided Attention Mass Allocation for Training-Free Multi-Style Transfer

论文配图:MAST: Mask-Guided Attention Mass Allocation for Training-Free Multi-Style Transfer
图 1 · 摘自论文原文
  • 用掩码引导注意力分配,精准控制内容与风格交互。
  • 多风格融合无边界伪影,结构一致性显著提升。
  • 适合需要快速多风格转换的图像生成场景。

风格迁移旨在保留内容图像的语义布局和结构几何的同时,赋予其参考风格的视觉特征。尽管基于扩散模型的方法凭借强大的生成先验和可控内部表示展现出优异的风格化能力,但通常假设单一全局风格。扩展至多风格场景时常引发边界伪影、风格不稳定和结构不一致问题,源于多种风格表示间的干扰。为此,我们提出MAST(Mask-Guided Attention Mass Allocation for Training-Free Multi-Style Transfer),一种无需训练的新型框架,通过显式调控扩散注意力机制内的内容-风格交互。MAST集成四个协同模块:首先,布局保持查询锚定通过内容查询稳固语义结构,防止全局布局坍塌;其次,对数级别注意力质量分配在空间区域间确定性地分配注意力概率质量,无缝融合多种风格且无边界伪影;第三,锐度感知温度调节恢复因多风格扩展导致的注意力锐度下降;最后,差异感知细节注入通过测量结构差异自适应补偿局部高频细节损失。大量实验表明,MAST有效缓解边界伪影并维持结构一致性,在增加应用风格数量时仍保持纹理保真度和空间连贯性。

原文摘要 · Abstract (English)

Style transfer aims to render a content image with the visual characteristics of a reference style while preserving its underlying semantic layout and structural geometry. While recent diffusion-based models demonstrate strong stylization capabilities by leveraging powerful generative priors and controllable internal representations, they typically assume a single global style. Extending them to multi-style scenarios often leads to boundary artifacts, unstable stylization, and structural inconsistency due to interference between multiple style representations. To overcome these limitations, we propose MAST (Mask-Guided Attention Mass Allocation for Training-Free Multi-Style Transfer), a novel training-free framework that explicitly controls content-style interactions within the diffusion attention mechanism. To achieve artifact-free and structure-preserving stylization, MAST integrates four connected modules. First, Layout-preserving Query Anchoring prevents global layout collapse by firmly anchoring the semantic structure using content queries. Second, Logit-level Attention Mass Allocation deterministically distributes attention probability mass across spatial regions, seamlessly fusing multiple styles without boundary artifacts. Third, Sharpness-aware Temperature Scaling restores the attention sharpness degraded by multi-style expansion. Finally, Discrepancy-aware Detail Injection adaptively compensates for localized high-frequency detail losses by measuring structural discrepancies. Extensive experiments demonstrate that MAST effectively mitigates boundary artifacts and maintains structural consistency, preserving texture fidelity and spatial coherence even as the number of applied styles increases.

风格迁移扩散模型多风格无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。