M2-omni实现多模态统一建模,支持任意组合输入输出。
M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
- 采用统一序列建模框架,整合音视频图文输入输出
- 在多个数据集上表现接近GPT-4o,保持纯文本任务强性能
- 提出动态平衡策略,解决多模态训练进度不一问题
我们提出M2-omni,一个前沿的开源全模态多模态大模型(omni-MLLM),其性能可媲美GPT-4o。M2-omni采用统一多模态序列建模框架,使大语言模型具备全面的跨模态理解与生成能力。具体而言,该模型可处理音频、视频、图像和文本的任意组合输入,并生成交错的音频、图像或文本输出,从而实现高级交互式实时体验。全模态大模型的训练面临各模态数据量差异大、收敛速度不一致的挑战。为此,我们提出预训练阶段的步数平衡策略以应对模态数据量差异;在指令微调阶段引入动态自适应平衡策略,同步各模态训练进度,确保最优收敛。特别地,我们优先保障纯文本任务的强性能,以维持模型语言理解能力的鲁棒性。据我们所知,M2-omni是当前最接近GPT-4o的开源模型之一,具有全面的模态与任务支持能力及卓越性能。我们期望M2-omni能推动全模态大模型的发展,促进该领域未来研究。
原文摘要 · Abstract (English)
We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which empowers Large Language Models(LLMs) to acquire comprehensive cross-modal understanding and generation capabilities. Specifically, M2-omni can process arbitrary combinations of audio, video, image, and text modalities as input, generating multimodal sequences interleaving with audio, image, or text outputs, thereby enabling an advanced and interactive real-time experience. The training of such an omni-MLLM is challenged by significant disparities in data quantity and convergence rates across modalities. To address these challenges, we propose a step balance strategy during pre-training to handle the quantity disparities in modality-specific data. Additionally, a dynamically adaptive balance strategy is introduced during the instruction tuning stage to synchronize the modality-wise training progress, ensuring optimal convergence. Notably, we prioritize preserving strong performance on pure text tasks to maintain the robustness of M2-omni's language understanding capability throughout the training process. To our best knowledge, M2-omni is currently a very competitive open-source model to GPT-4o, characterized by its comprehensive modality and task support, as well as its exceptional performance. We expect M2-omni will advance the development of omni-MLLMs, thus facilitating future research in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。