无需训练,用扩散模型实现场景文字自由风格迁移。
SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
- 通过特征注入与注意力机制,精准传递文本区域风格。
- 在保持文字可读性的同时,实现高保真局部风格转换。
- 适合需要快速样式调整的图像编辑与设计场景。
随着扩散模型的快速发展,风格迁移已取得显著进展。然而,场景文字的灵活且局部化的风格编辑仍是未解难题。现有方法虽能实现文本区域修改,但通常仅限于内容替换和简单风格,缺乏自由风格迁移能力。本文提出SceneTextStylizer,一种基于扩散模型的无需训练框架,实现场景图像中文本的灵活、高保真风格迁移。与以往仅进行全局风格转移或仅修改文本内容的方法不同,本方法支持提示词引导的文本区域风格变换,同时保持文字可读性和风格一致性。为此,我们设计了特征注入模块,结合扩散模型反演与自注意力机制,有效传递风格特征;引入基于距离的动态掩码机制,在每一步去噪过程中实现精确空间控制;并基于傅里叶变换构建风格增强模块,提升风格丰富性。大量实验表明,该方法在视觉保真度与文字保留方面均优于现有最先进方法。
原文摘要 · Abstract (English)
With the rapid development of diffusion models, style transfer has made remarkable progress. However, flexible and localized style editing for scene text remains an unsolved challenge. Although existing scene text editing methods have achieved text region editing, they are typically limited to content replacement and simple styles, which lack the ability of free-style transfer. In this paper, we introduce SceneTextStylizer, a novel training-free diffusion-based framework for flexible and high-fidelity style transfer of text in scene images. Unlike prior approaches that either perform global style transfer or focus solely on textual content modification, our method enables prompt-guided style transformation specifically for text regions, while preserving both text readability and stylistic consistency. To achieve this, we design a feature injection module that leverages diffusion model inversion and self-attention to transfer style features effectively. Additionally, a region control mechanism is introduced by applying a distance-based changing mask at each denoising step, enabling precise spatial control. To further enhance visual quality, we incorporate a style enhancement module based on the Fourier transform to reinforce stylistic richness. Extensive experiments demonstrate that our method achieves superior performance in scene text style transformation, outperforming existing state-of-the-art methods in both visual fidelity and text preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。