arXiv:2602.20903cs.CV2026-02中稿 · CVPR被引 11

让AI生成文字更准,自动识别并修复结构错误

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

  • 用可插拔的强化学习策略感知文字结构异常
  • 在中文文本上提升4%结构保真度和8.7%语义对齐
  • 适合需要高精度文字生成的研究与应用

视觉文本渲染(VTR)在文生图任务中仍是关键挑战,即使先进模型也常产生扭曲、模糊、错位等结构异常。我们发现主流多模态大模型和专业OCR模型大多无法感知这些异常,导致评估与强化学习优化受阻。为此提出TextPecker,一种即插即用的结构异常感知强化学习策略,适用于任意文生图生成器。通过构建字符级结构异常标注数据集,并开发笔画编辑合成引擎扩展异常覆盖范围。实验表明,TextPecker显著提升多种文生图模型性能;即使在已优化的Qwen-Image上,中文文本渲染仍实现平均4%结构保真度与8.7%语义对齐提升,创下新基准。本工作填补了高保真视觉文本生成的优化空白。

原文摘要 · Abstract (English)

Visual Text Rendering (VTR) remains a critical challenge in text-to-image generation, where even advanced models frequently produce text with structural anomalies such as distortion, blurriness, and misalignment. However, we find that leading MLLMs and specialist OCR models largely fail to perceive these structural anomalies, creating a critical bottleneck for both VTR evaluation and RL-based optimization. As a result, even state-of-the-art generators (e.g., Seedream4.0, Qwen-Image) still struggle to render structurally faithful text. To address this, we propose TextPecker, a plug-and-play structural anomaly perceptive RL strategy that mitigates noisy reward signals and works with any textto-image generator. To enable this capability, we construct a recognition dataset with character-level structural-anomaly annotations and develop a stroke-editing synthesis engine to expand structural-error coverage. Experiments show that TextPecker consistently improves diverse text-to-image models; even on the well-optimized Qwen-Image, it significantly yields average gains of 4% in structural fidelity and 8.7% in semantic alignment for Chinese text rendering, establishing a new state-of-the-art in high-fidelity VTR. Our work fills a gap in VTR optimization, providing a foundational step towards reliable and structural faithful visual text generation.

视觉文本文生图强化学习结构优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。