无需训练即可生成多语言Logo,保持文字结构清晰。
LogoDiffuser: Training-Free Multilingual Logo Generation and Stylization via Letter-Aware Attention Control
- 用字符图像替代文本提示,精准控制多语言文字结构。
- 通过核心注意力图注入,实现文字与视觉风格统一融合。
- 无需微调,支持任意语言,适合设计工具集成。
近年来文本到图像生成技术取得显著进展,但如何生成视觉与文字和谐统一的多语言设计类标志仍具挑战。现有方法在应用创意风格时常扭曲字符几何结构,且难以在不额外训练的前提下支持多语言文本生成。为此,我们提出LogoDiffuser,一种无需训练的方法,基于多模态扩散变换器合成多语言标志设计。不同于传统文本提示,我们直接输入目标字符的图像,从而无论何种语言都能稳健控制字符结构。通过分析联合注意力机制,识别出对文本结构响应强烈的“核心令牌”,并将其最具信息量的注意力图注入模型。进一步地,采用层间聚合注意力图,缓解不同层级间的注意力偏移,获得一致的核心令牌。大量实验与用户研究显示,该方法在多语言标志生成任务中达到当前最优性能。
原文摘要 · Abstract (English)
Recent advances in text-to-image generation have been remarkable, but generating multilingual design logos that harmoniously integrate visual and textual elements remains a challenging task. Existing methods often distort character geometry when applying creative styles and struggle to support multilingual text generation without additional training. To address these challenges, we propose LogoDiffuser, a training-free method that synthesizes multilingual logo designs using the multimodal diffusion transformer. Instead of using textual prompts, we input the target characters as images, enabling robust character structure control regardless of language. We first analyze the joint attention mechanism to identify core tokens, which are tokens that strongly respond to textual structures. With this observation, our method integrates character structure and visual design by injecting the most informative attention maps. Furthermore, we perform layer-wise aggregation of attention maps to mitigate attention shifts across layers and obtain consistent core tokens. Extensive experiments and user studies demonstrate that our method achieves state-of-the-art performance in multilingual logo generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。