Show-o2统一建模图文视频,实现跨模态生成与理解。
Show-o2: Improved Native Unified Multimodal Models

- 基于3D因果变分自编码空间,双路径融合构建统一视觉表征。
- 采用两阶段训练,支持更大模型规模,提升跨模态任务性能。
- 适合需要多模态生成与理解的开发者,开源可复用。
本文提出改进的原生统一多模态模型 Show-o2,采用自回归建模与流匹配技术。基于3D因果变分自编码器空间,通过空间(时间)双路径融合构建统一视觉表征,实现图像与视频模态的可扩展性,保障多模态理解与生成效果。在语言模型基础上,自回归建模应用于语言头,流匹配应用于流头,分别实现文本词元预测与图像/视频生成。设计两阶段训练方案,有效学习并扩展至更大模型。结果表明,Show-o2 在多种模态(文本、图像、视频)的多模态理解与生成任务中表现出强大通用性。代码与模型已开源于 https://github.com/showlab/Show-o。
原文摘要 · Abstract (English)
This paper presents improved native unified multimodal models, \emph{i.e.,} Show-o2, that leverage autoregressive modeling and flow matching. Built upon a 3D causal variational autoencoder space, unified visual representations are constructed through a dual-path of spatial (-temporal) fusion, enabling scalability across image and video modalities while ensuring effective multimodal understanding and generation. Based on a language model, autoregressive modeling and flow matching are natively applied to the language head and flow head, respectively, to facilitate text token prediction and image/video generation. A two-stage training recipe is designed to effectively learn and scale to larger models. The resulting Show-o2 models demonstrate versatility in handling a wide range of multimodal understanding and generation tasks across diverse modalities, including text, images, and videos. Code and models are released at https://github.com/showlab/Show-o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。