用智能工作流提升复杂文字与公式生成精度
GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
- 通过代理式流程引入字形模板,优化潜在空间与注意力图
- 无需训练即可适配多种文生图模型,生成更精准
- 专为复杂字符与公式设计基准,适合需要高精度文本渲染的场景
尽管生成模型在文本渲染方面取得显著进展,但准确生成复杂文字和数学公式仍是重大挑战,主要源于现有模型在面对分布外提示时指令遵循能力有限。为此,我们提出GlyphBanana及相应基准,专门用于复杂字符与公式的渲染。GlyphBanana采用代理式工作流,结合辅助工具将字形模板注入潜在空间和注意力图中,实现生成图像的迭代优化。其无训练方法可无缝适配多种文生图(T2I)模型,在精度上优于现有基线。大量实验验证了该工作流的有效性。相关代码已公开于https://github.com/yuriYanZeXuan/GlyphBanana。
原文摘要 · Abstract (English)
Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable challenge. This difficulty primarily stems from the limited instruction-following capabilities of current models when encountering out-of-distribution prompts. To address this, we introduce GlyphBanana, alongside a corresponding benchmark specifically designed for rendering complex characters and formulas. GlyphBanana employs an agentic workflow that integrates auxiliary tools to inject glyph templates into both the latent space and attention maps, facilitating the iterative refinement of generated images. Notably, our training-free approach can be seamlessly applied to various Text-to-Image (T2I) models, achieving superior precision compared to existing baselines. Extensive experiments demonstrate the effectiveness of our proposed workflow. Associated code is publicly available at https://github.com/yuriYanZeXuan/GlyphBanana.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。