提出ILLUME-X模型,实现高质量自由混合图文生成。
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

- 构建优化的数据管道与渐进式训练策略,提升图文交错生成能力。
- 在风格迁移、图像分解等任务上超越现有统一模型表现。
- 适合需要灵活图文生成的创作类应用开发者使用。
生成式AI模型在文本与图像联合生成方面取得重要进展,尤其适用于模态交错的任务。为推动该领域迈向新阶段,模型需能自主生成自由形式的交错图文序列。本文提出ILLUME-X,一种先进的统一多模态范式,通过提升多模态数据效率和稳定训练过程,实现高质量自由交错图文生成。ILLUME-X包含三个关键组件:(i) 针对交错图文生成优化的扩展训练数据流水线;(ii) 采用自适应目标的渐进式训练策略,支持自由长度多模态标记序列;(iii) 用于交错图文序列的客观且全面的评估方法ILScore。实验表明,ILLUME-X在风格迁移、图像分解和故事生成等任务中均优于先前统一模型。
原文摘要 · Abstract (English)
The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving the interleaving of both modalities. To advance this intelligence to the next stage, it is crucial for models to autonomously generate free-form interleaved text-image sequences. In this paper, we introduce ILLUME-X, an advanced unified multimodal paradigm that enables high-quality, free-form interleaved text-image generation by improving multimodal data efficiency and stabilizing the multimodal training process. ILLUME-X comprises three key components: (i) an expanded training data pipeline optimized for interleaved text-image generation, (ii) a progressive training strategy with self-adaptive objectives for free-length multimodal token sequences, and (iii) an objective and comprehensive evaluation method ILScore for interleaved text-image sequences. Notably, our ILLUME-X outperforms previous unified models across multiple interleaved text-image generation tasks like style transfer, image decomposition and storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。