用高质量数据提升文本生成图像的准确性和美感,效果显著。
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
- 基于Deepseek-R1构建1024×1024高清图文数据集LeX-10K
- LeX-Lumina在CreateBench上文本准确率提升79.81%(PNED)
- 适合关注文本渲染精度与视觉美学的研究者和开发者
我们提出LeX-Art,一套以数据为中心的高质量文本图像合成方案,系统性弥合提示表达力与文本呈现保真度之间的差距。该方法基于Deepseek-R1构建数据合成流水线,创建包含10,000张高分辨率、美学优化的1024×1024图像数据集LeX-10K。除数据构建外,还开发了鲁棒的提示增强模型LeX-Enhancer,训练出两个文生图模型LeX-FLUX与LeX-Lumina,实现顶尖的文本渲染性能。为系统评估视觉文本生成,提出LeX-Bench基准,涵盖保真度、美学和对齐度,并引入新型指标Pairwise Normalized Edit Distance(PNED),用于鲁棒的文本准确性评估。实验表明,LeX-Lumina在CreateBench上获得79.81%的PNED提升,LeX-FLUX在颜色(+3.18%)、位置(+4.45%)和字体准确率(+3.81%)上优于基线。代码、模型、数据集及演示均已开源。
原文摘要 · Abstract (English)
We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a high-quality data synthesis pipeline based on Deepseek-R1 to curate LeX-10K, a dataset of 10K high-resolution, aesthetically refined 1024$\times$1024 images. Beyond dataset construction, we develop LeX-Enhancer, a robust prompt enrichment model, and train two text-to-image models, LeX-FLUX and LeX-Lumina, achieving state-of-the-art text rendering performance. To systematically evaluate visual text generation, we introduce LeX-Bench, a benchmark that assesses fidelity, aesthetics, and alignment, complemented by Pairwise Normalized Edit Distance (PNED), a novel metric for robust text accuracy evaluation. Experiments demonstrate significant improvements, with LeX-Lumina achieving a 79.81% PNED gain on CreateBench, and LeX-FLUX outperforming baselines in color (+3.18%), positional (+4.45%), and font accuracy (+3.81%). Our codes, models, datasets, and demo are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。