arXiv:2409.17692cs.CLcs.AI2024-09EMNLP被引 29

MIO能端到端生成多模态内容,支持跨模态自由交互。

MIO: A Foundation Model on Multimodal Tokens

  • 用统一离散令牌构建多模态模型,实现语音、文本、图像、视频的联合建模。
  • 在多项任务上超越双模态基线,部分能力超过专用模型。
  • 适合需要跨模态生成与推理的研究者,如视频-文本协同创作。

本文提出MIO,一种基于多模态令牌的新型基础模型,可端到端、自回归地理解与生成语音、文本、图像和视频。尽管大语言模型(LLMs)和多模态大语言模型(MM-LLMs)推动了通用人工智能的发展,但依然缺乏真正的任意模态间理解与生成能力。GPT-4o虽展示任意模态交互潜力,但为闭源且不支持多模态交错序列生成。为此,我们提出MIO,通过因果多模态建模,在四种模态的离散令牌混合数据上训练,经历四阶段:对齐预训练、交错预训练、语音增强预训练及多样化文本、视觉、语音任务的监督微调。实验表明,MIO性能与先前双模态基线相当甚至更优,部分任务超越任何模态基线及专用模型。此外,其任意模态能力展现出交错视频-文本生成、视觉链式思维推理、视觉指引生成、指令式图像编辑等先进功能。

原文摘要 · Abstract (English)

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language models (LLMs) and multimodal large language models (MM-LLMs) propels advancements in artificial general intelligence through their versatile capabilities, they still lack true any-to-any understanding and generation. Recently, the release of GPT-4o has showcased the remarkable potential of any-to-any LLMs for complex real-world tasks, enabling omnidirectional input and output across images, speech, and text. However, it is closed-source and does not support the generation of multimodal interleaved sequences. To address this gap, we present MIO, which is trained on a mixture of discrete tokens across four modalities using causal multimodal modeling. MIO undergoes a four-stage training process: (1) alignment pre-training, (2) interleaved pre-training, (3) speech-enhanced pre-training, and (4) comprehensive supervised fine-tuning on diverse textual, visual, and speech tasks. Our experimental results indicate that MIO exhibits competitive, and in some cases superior, performance compared to previous dual-modal baselines, any-to-any model baselines, and even modality-specific baselines. Moreover, MIO demonstrates advanced capabilities inherent to its any-to-any feature, such as interleaved video-text generation, chain-of-visual-thought reasoning, visual guideline generation, instructional image editing, etc.

多模态生成模型基础模型自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。