构建文本复杂度驱动的图文生成评测基准,揭示现有模型在文字嵌入中的关键缺陷。
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
- 设计多维度文本与提示复杂度的评测集,覆盖多种文字特征和场景
- 发现扩散模型普遍存在拼写错误和语义不一致问题,影响图文一致性
- 适用于评估视觉文本生成模型,尤其适合关注多模态内容生成的研究者
在教育材料、广告等多模态视觉内容自动生成中,将文字准确嵌入图像至关重要。然而,现有的基于扩散模型的文本到图像生成方法常在拼写准确性、上下文相关性和视觉连贯性方面表现不佳。由于缺乏全面的评测基准,评估这些模型的文字嵌入能力变得困难。本文提出 TextInVision,一个大规模、以文本和提示复杂度为导向的评测基准,用于评估扩散模型在图像中有效集成视觉文本的能力。我们设计了多样化的提示和文本,涵盖不同属性与文字特征。同时构建了图像数据集,用于测试变分自编码器(VAE)模型在不同字符表示下的表现,表明VAE架构本身也会在扩散框架中带来文本生成挑战。通过对多个模型的广泛分析,我们识别出常见错误,如拼写错误和上下文错配。通过定位不同提示和文本下的失败点,本研究为未来AI生成多模态内容的技术进步奠定了基础。
原文摘要 · Abstract (English)
Generating images with embedded text is crucial for the automatic production of visual and multimodal documents, such as educational materials and advertisements. However, existing diffusion-based text-to-image models often struggle to accurately embed text within images, facing challenges in spelling accuracy, contextual relevance, and visual coherence. Evaluating the ability of such models to embed text within a generated image is complicated due to the lack of comprehensive benchmarks. In this work, we introduce TextInVision, a large-scale, text and prompt complexity driven benchmark designed to evaluate the ability of diffusion models to effectively integrate visual text into images. We crafted a diverse set of prompts and texts that consider various attributes and text characteristics. Additionally, we prepared an image dataset to test Variational Autoencoder (VAE) models across different character representations, highlighting that VAE architectures can also pose challenges in text generation within diffusion frameworks. Through extensive analysis of multiple models, we identify common errors and highlight issues such as spelling inaccuracies and contextual mismatches. By pinpointing the failure points across different prompts and texts, our research lays the foundation for future advancements in AI-generated multimodal content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。