Ola模型在图文音多模态理解上达到顶尖水平,推动通用多模态模型发展。
Ola: Pushing the Frontiers of Omni-Modal Language Model
- 以视频为桥梁,分阶段对齐不同模态,提升跨模态理解能力。
- 在图像、视频、音频任务上均超越现有开源多模态模型,媲美专用模型性能。
- 开源全部代码、权重与数据,适合多模态研究与应用开发者使用。
GPT-4o之后,通用多模态大模型的发展备受关注。尽管已有开源替代方案出现,但在性能上仍落后于专用单模态模型。本文提出Ola,一个在图像、视频和音频理解上表现优异的通用多模态语言模型,显著推进了该领域的边界。我们系统探索了架构设计、数据构建与训练策略。Ola通过多项改进,增强了视觉理解与语音识别能力。同时,重新思考多模态训练中的跨模态关系,强调以视频为关键桥梁,并提出渐进式训练流程:从差异最大的模态开始,逐步过渡到更接近的模态对齐。大量实验表明,Ola在所有模态上均优于现有开源多模态模型,且与同规模顶尖专用模型相比表现相当。我们致力于将Ola打造为全开源的多模态理解解决方案,以促进该领域未来发展。模型权重、代码及数据已公开于https://github.com/Ola-Omni/Ola。
原文摘要 · Abstract (English)
Recent advances in large language models, particularly following GPT-4o, have sparked increasing interest in developing omni-modal models capable of understanding more modalities. While some open-source alternatives have emerged, there is still a notable lag behind specialized single-modality models in performance. In this paper, we present Ola, an Omni-modal Language model that achieves competitive performance across image, video, and audio understanding compared to specialized counterparts, pushing the frontiers of the omni-modal language model to a large extent. We conduct a comprehensive exploration of architectural design, data curation, and training strategies essential for building a robust omni-modal model. Ola incorporates advanced visual understanding and audio recognition capabilities through several critical and effective improvements over mainstream baselines. Moreover, we rethink inter-modal relationships during omni-modal training, emphasizing cross-modal alignment with video as a central bridge, and propose a progressive training pipeline that begins with the most distinct modalities and gradually moves towards closer modality alignment. Extensive experiments demonstrate that Ola surpasses existing open omni-modal LLMs across all modalities while achieving highly competitive performance compared to state-of-the-art specialized models of similar sizes. We aim to make Ola a fully open omni-modal understanding solution to advance future research in this emerging field. Model weights, code, and data are open-sourced at https://github.com/Ola-Omni/Ola.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。