arXiv:2505.18985cs.LGcs.CL2025-05EMNLP被引 10

测试扩散模型生成图像中文本的准确与一致性,发现其存在长距离依赖缺陷。

STRICT: Stress Test of Rendering Images Containing Text

  • 构建多维度基准测试,评估文本长度、可读性与指令遵循度。
  • 主流模型在超过10个字符的文本生成中错误率超30%。
  • 适合研究图文生成、模型评估与提示工程的学者参考。

尽管扩散模型在文本到图像生成中实现了逼真且多样场景的合成,但在图像中生成一致且可读的文本方面仍存在困难。这一不足通常归因于基于扩散的生成固有的局部偏差,限制了其对长程空间依赖的建模能力。本文提出$ extbf{STRICT}$,一个系统化的基准测试,用于评估扩散模型在图像中生成连贯且符合指令的文本的能力。该基准从三个维度进行评估:(1)可生成文本的最大长度;(2)生成文本的正确性与可读性;(3)未遵循指令生成文本的比例。我们测试了多个前沿模型,包括专有与开源版本,揭示了在长程一致性与指令遵循方面的持续局限。研究结果为架构瓶颈提供了洞见,并推动了多模态生成建模的未来方向。完整评估流程已发布于https://github.com/tianyu-z/STRICT-Bench。

原文摘要 · Abstract (English)

While diffusion models have revolutionized text-to-image generation with their ability to synthesize realistic and diverse scenes, they continue to struggle to generate consistent and legible text within images. This shortcoming is commonly attributed to the locality bias inherent in diffusion-based generation, which limits their ability to model long-range spatial dependencies. In this paper, we introduce $\textbf{STRICT}$, a benchmark designed to systematically stress-test the ability of diffusion models to render coherent and instruction-aligned text in images. Our benchmark evaluates models across multiple dimensions: (1) the maximum length of readable text that can be generated; (2) the correctness and legibility of the generated text, and (3) the ratio of not following instructions for generating text. We evaluate several state-of-the-art models, including proprietary and open-source variants, and reveal persistent limitations in long-range consistency and instruction-following capabilities. Our findings provide insights into architectural bottlenecks and motivate future research directions in multimodal generative modeling. We release our entire evaluation pipeline at https://github.com/tianyu-z/STRICT-Bench.

文本生成扩散模型图像评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。