arXiv:2601.06605cs.CV2026-01

无需训练,用参考图实现精准风格迁移。

Sissi: Zero-shot Style-guided Image Synthesis via Semantic-style Integration

  • 将风格生成转为上下文学习任务,融合语义与视觉提示。
  • 在COCO、Cityscapes上实现高保真风格迁移,平衡内容与风格。
  • 适合快速风格化应用,无需微调或复杂计算。

文本引导的图像生成已因大规模扩散模型取得显著进展,但借助视觉样例实现精确风格化仍具挑战。现有方法常依赖特定任务重训练或昂贵的反演过程,易损害内容完整性、降低风格保真度,并在语义遵循与风格对齐间形成不良权衡。本文提出一种无需训练的框架,将风格引导生成重构为上下文学习任务。通过文本语义提示引导,将参考风格图与掩码目标图拼接,利用预训练的ReFlow inpainting模型,借助多模态注意力融合,实现语义内容与目标风格的无缝整合。进一步分析多模态注意力融合中的不平衡与噪声敏感性问题,提出动态语义-风格融合(DSSI)机制,重新加权文本语义与视觉特征之间的注意力,有效缓解指导冲突,提升输出一致性。实验表明,该方法在保持语义-风格平衡的同时,实现了高保真风格化与优异视觉质量,为复杂且易出错的先前方法提供了一种简洁而强大的替代方案。

原文摘要 · Abstract (English)

Text-guided image generation has advanced rapidly with large-scale diffusion models, yet achieving precise stylization with visual exemplars remains difficult. Existing approaches often depend on task-specific retraining or expensive inversion procedures, which can compromise content integrity, reduce style fidelity, and lead to an unsatisfactory trade-off between semantic prompt adherence and style alignment. In this work, we introduce a training-free framework that reformulates style-guided synthesis as an in-context learning task. Guided by textual semantic prompts, our method concatenates a reference style image with a masked target image, leveraging a pretrained ReFlow-based inpainting model to seamlessly integrate semantic content with the desired style through multimodal attention fusion. We further analyze the imbalance and noise sensitivity inherent in multimodal attention fusion and propose a Dynamic Semantic-Style Integration (DSSI) mechanism that reweights attention between textual semantic and style visual tokens, effectively resolving guidance conflicts and enhancing output coherence. Experiments show that our approach achieves high-fidelity stylization with superior semantic-style balance and visual quality, offering a simple yet powerful alternative to complex, artifact-prone prior methods.

图像生成风格迁移扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。