arXiv:2507.09573cs.CV2025-07

让文字生成支持局部修改与多轮优化,还能理解抽象指令。

WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending

  • 用注意力机制实现多区域精准控制,无需重新训练。
  • 通过噪声融合技术实现连续迭代,保持图像质量不下降。
  • 结合大模型解析模糊指令,适合设计师快速试错。

艺术字体旨在对输入字符进行兼具创意与可读性的视觉风格化处理。传统方法依赖手工设计,而基于扩散模型的生成方法虽实现了自动化,但仍受限于交互性不足,缺乏对局部修改、迭代优化、多字符组合及开放式提示的理解能力。我们提出 WordCraft,一个集成扩散模型的交互式艺术字体系统。该系统采用无训练的区域注意力机制,实现多区域精确生成;引入噪声融合策略,支持持续优化而不损失视觉质量;并整合大语言模型,解析具体与抽象用户提示,实现意图驱动的灵活生成。系统可在单字与多字输入下跨多种语言生成高质量风格化字体,支持多样化的用户工作流。本研究显著提升了艺术字体生成的交互性,为创作者提供了更多可能性。

原文摘要 · Abstract (English)

Artistic typography aims to stylize input characters with visual effects that are both creative and legible. Traditional approaches rely heavily on manual design, while recent generative models, particularly diffusion-based methods, have enabled automated character stylization. However, existing solutions remain limited in interactivity, lacking support for localized edits, iterative refinement, multi-character composition, and open-ended prompt interpretation. We introduce WordCraft, an interactive artistic typography system that integrates diffusion models to address these limitations. WordCraft features a training-free regional attention mechanism for precise, multi-region generation and a noise blending that supports continuous refinement without compromising visual quality. To support flexible, intent-driven generation, we incorporate a large language model to parse and structure both concrete and abstract user prompts. These components allow our framework to synthesize high-quality, stylized typography across single- and multi-character inputs across multiple languages, supporting diverse user-centered workflows. Our system significantly enhances interactivity in artistic typography synthesis, opening up creative possibilities for artists and designers.

艺术字体扩散模型交互生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。