用1维文本感知标记器,让开源数据训练出媲美私有数据的文生图模型。
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
- 设计1维文本感知标记器,解码时融合文本信息提升效率
- 单阶段训练替代复杂两阶段蒸馏,可扩展至大规模数据集
- 基于开源数据训练出性能媲美私有数据的文生图模型
图像标记器是现代文生图生成模型的基础,但训练困难。现有大多数文生图模型依赖大规模高质量私有数据集,难以复现。本文提出文本感知的1维标记器(TA-TiTok),可使用离散或连续1维标记,独特地在解码阶段融入文本信息,加速收敛并提升性能。其简化的一阶段训练过程取代了以往1维标记器复杂的两阶段蒸馏,实现对大规模数据集的无缝扩展。基于此,我们构建了一类仅在开放数据上训练的文生图掩码生成模型(MaskGen),性能可媲美基于私有数据训练的模型。我们计划发布高效强大的TA-TiTok标记器及开源权重的MaskGen模型,推动文生图掩码生成模型的普及与民主化。
原文摘要 · Abstract (English)
Image tokenizers form the foundation of modern text-to-image generative models but are notoriously difficult to train. Furthermore, most existing text-to-image models rely on large-scale, high-quality private datasets, making them challenging to replicate. In this work, we introduce Text-Aware Transformer-based 1-Dimensional Tokenizer (TA-TiTok), an efficient and powerful image tokenizer that can utilize either discrete or continuous 1-dimensional tokens. TA-TiTok uniquely integrates textual information during the tokenizer decoding stage (i.e., de-tokenization), accelerating convergence and enhancing performance. TA-TiTok also benefits from a simplified, yet effective, one-stage training process, eliminating the need for the complex two-stage distillation used in previous 1-dimensional tokenizers. This design allows for seamless scalability to large datasets. Building on this, we introduce a family of text-to-image Masked Generative Models (MaskGen), trained exclusively on open data while achieving comparable performance to models trained on private data. We aim to release both the efficient, strong TA-TiTok tokenizers and the open-data, open-weight MaskGen models to promote broader access and democratize the field of text-to-image masked generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。