首个专注长文本图像生成的模型,解决复杂文档排版难题。
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
- 用专为文本优化的二值化分词器提升长文本生成质量。
- 在长文本图像生成上显著优于SD3.5 Large、DALL-E 3和GPT4o。
- 支持字体、颜色、对齐等属性灵活控制,适合文档自动化生成。
自回归与扩散模型在短文本图像生成方面已取得优异表现,但生成如幻灯片或文档中的段落级长文本仍面临重大挑战。本文首次聚焦长文本图像生成,填补现有文本到图像系统仅处理短语或单句的空白。通过对主流自回归模型的全面分析,我们识别出图像分词器是影响文本生成质量的关键瓶颈。为此,提出一种专为场景文本设计的新型文本导向二值分词器。基于该分词器,构建了多模态自回归模型 exttt{ModelName},在长文本图像生成中实现前所未有的保真度。模型具备强可控性,可灵活定制字体样式、大小、颜色及对齐方式。大量实验表明, exttt{ModelName}在准确率、一致性和灵活性上显著超越SD3.5 Large、DALL-E 3和GPT4o。除技术突破外,该模型还为文档与演示文稿的混合生成等创新应用开辟新可能,树立长文本图像生成新范式。
原文摘要 · Abstract (English)
Recent advancements in autoregressive and diffusion models have led to strong performance in image generation with short scene text words. However, generating coherent, long-form text in images, such as paragraphs in slides or documents, remains a major challenge for current generative models. We present the first work specifically focused on long text image generation, addressing a critical gap in existing text-to-image systems that typically handle only brief phrases or single sentences. Through comprehensive analysis of state-of-the-art autoregressive generation models, we identify the image tokenizer as a critical bottleneck in text generating quality. To address this, we introduce a novel text-focused, binary tokenizer optimized for capturing detailed scene text features. Leveraging our tokenizer, we develop \ModelName, a multimodal autoregressive model that excels in generating high-quality long-text images with unprecedented fidelity. Our model offers robust controllability, enabling customization of text properties such as font style, size, color, and alignment. Extensive experiments demonstrate that \ModelName~significantly outperforms SD3.5 Large~\cite{sd3} and GPT4o~\cite{gpt4o} with DALL-E 3~\cite{dalle3} in generating long text accurately, consistently, and flexibly. Beyond its technical achievements, \ModelName~opens up exciting opportunities for innovative applications like interleaved document and PowerPoint generation, establishing a new frontier in long-text image generating.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。