为生成图像中的文字准确性提供新评估方法
ABHINAW: A method for Automatic Evaluation of Typography within AI-Generated Images
- 基于逐字匹配计算文本准确率,处理重复、大小写等干扰
- 提出简洁性调整机制,解决生成文字过长问题
- 可帮助设计师和研究者优化AI图像文字生成效果
在生成式AI快速发展的背景下,MidJourney、DALL-E和Stable Diffusion等平台虽能生成高质量图像,但常无法准确生成图像中的文字。若能实现零样本下文字生成的高精度,将使AI图像更具意义,并推动图形设计民主化。首要挑战是建立可靠的文本与排版评估体系。现有基准如CLIP SCORE和T2I-CompBench++未能系统评估扩散模型生成图像中的文字表现。本文提出一种专用于评估AI生成图像中文字与排版性能的新评分矩阵。采用逐字匹配策略,从参考文本到生成文本进行精确匹配评分,有效处理重复词、大小写敏感、词语混杂、字符错位等冗余问题。同时提出新颖的‘简洁性调整’方法,以应对生成内容过长的问题。还对高频词与低频词引发的常见错误进行了定量分析。
原文摘要 · Abstract (English)
In the fast-evolving field of Generative AI, platforms like MidJourney, DALL-E, and Stable Diffusion have transformed Text-to-Image (T2I) Generation. However, despite their impressive ability to create high-quality images, they often struggle to generate accurate text within these images. Theoretically, if we could achieve accurate text generation in AI images in a ``zero-shot'' manner, it would not only make AI-generated images more meaningful but also democratize the graphic design industry. The first step towards this goal is to create a robust scoring matrix for evaluating text accuracy in AI-generated images. Although there are existing bench-marking methods like CLIP SCORE and T2I-CompBench++, there's still a gap in systematically evaluating text and typography in AI-generated images, especially with diffusion-based methods. In this paper, we introduce a novel evaluation matrix designed explicitly for quantifying the performance of text and typography generation within AI-generated images. We have used letter by letter matching strategy to compute the exact matching scores from the reference text to the AI generated text. Our novel approach to calculate the score takes care of multiple redundancies such as repetition of words, case sensitivity, mixing of words, irregular incorporation of letters etc. Moreover, we have developed a Novel method named as brevity adjustment to handle excess text. In addition we have also done a quantitative analysis of frequent errors arise due to frequently used words and less frequently used words. Project page is available at: https://github.com/Abhinaw3906/ABHINAW-MATRIX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。