arXiv:2609.02349cs.CV2026-09

通过锚定字形位置提升图像中文本渲染精度。

GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors

论文配图:GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors
图 1 · 摘自论文原文
  • 用轻量级字形补丁条件,结合位置编码锚定目标图像。
  • 在长文本、复杂排版和罕见字符场景下显著提升文本保真度。
  • 适合需要高精度文本生成与编辑的视觉模型研究者。

图像生成与编辑模型在渲染准确文本方面仍面临挑战,尤其在长篇、复杂且密集排列的文本或罕见字符场景下。现有方法或依赖更强的主干网络和数据驱动训练,缺乏显式字形先验;或通过专用设计引入字形先验,但鲁棒性不足。本文提出GlyphAnchor,一种针对文生图与图像编辑扩散Transformer模型的文本渲染增强方法。该方法通过将轻量级字形补丁条件的位置锚定于目标图像(利用模型原生位置编码),增强主干网络表现力。采用分阶段监督微调训练,并通过文本感知后训练进一步提升鲁棒性。同时提出InfoTextBench,用于评估生成与编辑场景下的文本丰富视觉渲染效果。在多个主干网络与基准测试中,包括长文本、复杂布局及罕见字符场景,GlyphAnchor均持续提升文本保真度,同时保持整体图像质量。

原文摘要 · Abstract (English)

Rendering accurate text remains difficult for image generation and editing models, especially when the target contains long, complex, and densely arranged text or rare characters. Existing approaches either improve native text rendering through stronger backbones and data-centric training without explicit glyph priors, or incorporate glyph priors through specialized designs that remain insufficiently accurate and robust under challenging scenarios. We introduce GlyphAnchor, a novel text-rendering enhancement method for both text-to-image and image-editing diffusion transformer models. GlyphAnchor enhances the backbone with lightweight glyph patch conditions whose positions are anchored to the target image through the model's native positional encoding. We train this capability with staged supervised finetuning and further refine it with text-aware post-training to improve robustness. We also introduce InfoTextBench, a benchmark for evaluating text-rich visual text rendering in both generation and editing settings. Experiments across multiple backbones and benchmarks, including long, complex, and densely arranged text and rare character scenarios, show that GlyphAnchor consistently improves text fidelity while preserving overall image quality.

文本生成扩散模型字形先验图像编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。