一个统一模型同时处理视频理解与生成,效率更高。
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- 用多模态预热策略提升模型能力,减少训练成本。
- 在图像和视频任务上表现优于同类先进模型。
- 适合需要高效统一处理多模态内容的研究者使用。
随着语言模型的发展,统一的多模态理解和生成取得了显著进展,模型架构从分离组件演变为统一的单模型框架。本文探索了一种高效的训练范式,构建一个用于统一多模态理解和生成的单一Transformer。具体地,我们提出一种利用先验知识的多模态预热策略以扩展模型能力。为解决跨模态兼容性问题,引入特征预缩放和多模态AdaLN技术。集成上述技术后,我们提出了HaploOmni——一种新型的单个多模态Transformer。在有限训练成本下,HaploOmni在多个图像和视频理解和生成基准测试中达到与先进统一模型相当的性能。代码将公开于https://github.com/Tencent/HaploVLM。
原文摘要 · Abstract (English)
With the advancement of language models, unified multimodal understanding and generation have made significant strides, with model architectures evolving from separated components to unified single-model frameworks. This paper explores an efficient training paradigm to build a single transformer for unified multimodal understanding and generation. Specifically, we propose a multimodal warmup strategy utilizing prior knowledge to extend capabilities. To address cross-modal compatibility challenges, we introduce feature pre-scaling and multimodal AdaLN techniques. Integrating the proposed technologies, we present the HaploOmni, a new single multimodal transformer. With limited training costs, HaploOmni achieves competitive performance across multiple image and video understanding and generation benchmarks over advanced unified models. All codes will be made public at https://github.com/Tencent/HaploVLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。