通过嵌入重构实现文本与图像风格对齐,提升风格迁移的多样性与控制力。
ArtCrafter: Text-Image Aligning Style Transfer via Embedding Reframing
- 引入注意力机制提取图像细粒度风格特征,增强表达精度。
- 构建跨模态对齐模块,使图文嵌入在共享空间中高效融合。
- 采用嵌入重构设计,显著提升生成结果的风格强度与多样性。
近年来,基于扩散模型的文本引导风格迁移取得显著进展,其凭借条件引导能力可有效控制生成过程。然而,直接的条件引导方法在平衡文本语义表达与输出多样性、捕捉风格特征方面仍面临挑战。为此,我们提出 ArtCrafter,一种新型文本到图像风格迁移框架。首先,设计基于注意力的风格提取模块,通过多层结构与 Perceiver 注意力机制,精准捕捉图像中的细微风格元素。其次,提出新颖的文本-图像对齐增强组件,通过注意力操作实现模态间平滑信息流动,有效平衡两种模态的控制力,将图像与文本嵌入映射至共享特征空间。最后,引入显式调制机制,通过嵌入重构设计,将多模态增强嵌入与原始嵌入无缝融合,从而支持多样化输出。大量实验表明,ArtCrafter 在视觉风格化任务中表现优异,展现出卓越的风格强度、可控性与多样性。
原文摘要 · Abstract (English)
Recent years have witnessed significant advancements in text-guided style transfer, primarily attributed to innovations in diffusion models. These models excel in conditional guidance, utilizing text or images to direct the sampling process. However, despite their capabilities, direct conditional guidance approaches often face challenges in balancing the expressiveness of textual semantics with the diversity of output results while capturing stylistic features. To address these challenges, we introduce ArtCrafter, a novel framework for text-to-image style transfer. Specifically, we introduce an attention-based style extraction module, meticulously engineered to capture the subtle stylistic elements within an image. This module features a multi-layer architecture that leverages the capabilities of perceiver attention mechanisms to integrate fine-grained information. Additionally, we present a novel text-image aligning augmentation component that adeptly balances control over both modalities, enabling the model to efficiently map image and text embeddings into a shared feature space. We achieve this through attention operations that enable smooth information flow between modalities. Lastly, we incorporate an explicit modulation that seamlessly blends multimodal enhanced embeddings with original embeddings through an embedding reframing design, empowering the model to generate diverse outputs. Extensive experiments demonstrate that ArtCrafter yields impressive results in visual stylization, exhibiting exceptional levels of stylistic intensity, controllability, and diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。