开源7B多模态大模型,支持图文音视频实时交互
Baichuan-Omni Technical Report
- 7B模型分两阶段对齐多模态并进行多任务微调
- 在多模态评测中表现优异,支持图像、视频、音频与文本联合处理
- 为开源社区提供强性能基准,适合多模态研究与应用开发
GPT-4o展现出突出的多模态能力与交互体验,但在开源领域尚无高性能对应模型。本文提出首个7B规模的开源多模态大语言模型Baichuan-Omni,具备同时处理图像、视频、音频与文本的能力,实现先进的多模态交互体验和强劲性能。通过从7B模型出发,采用两阶段训练策略:先进行多模态对齐,再在音频、图像、视频与文本上进行多任务微调,使语言模型有效理解视觉与听觉信息。在多个全模态与多模态基准测试中表现优异,旨在为开源社区提供强有力的性能基准,推动多模态理解与实时交互技术发展。
原文摘要 · Abstract (English)
The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart. In this paper, we introduce Baichuan-omni, the first open-source 7B Multimodal Large Language Model (MLLM) adept at concurrently processing and analyzing modalities of image, video, audio, and text, while delivering an advanced multimodal interactive experience and strong performance. We propose an effective multimodal training schema starting with 7B model and proceeding through two stages of multimodal alignment and multitask fine-tuning across audio, image, video, and text modal. This approach equips the language model with the ability to handle visual and audio data effectively. Demonstrating strong performance across various omni-modal and multimodal benchmarks, we aim for this contribution to serve as a competitive baseline for the open-source community in advancing multimodal understanding and real-time interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。