arXiv:2504.12315cs.CLcs.AI2025-04被引 1

轻量高效构建多模态大模型,支持文本图像视频音频理解

Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models

  • 分步构建框架,优化数据与训练流程提升效率
  • 在同类规模模型中多模态基准表现具竞争力
  • 开源模型权重、数据与代码,支持实时对话应用

随着多模态大语言模型(MLLMs)的发展,开源社区涌现出诸多优秀成果。然而,由于构建和训练多模态数据对非常复杂,打造强大MLLM仍需大量计算与时间成本。本文提出Capybara-OMNI,一种轻量高效训练的MLLM,支持文本、图像、视频和音频模态的理解。我们详细阐述了框架设计、数据构建与训练策略,逐步实现高性能模型。同时提供专属评测基准,用于验证跨模态理解能力。实验表明,遵循本方案可高效构建在各类多模态基准上表现优异的模型。为进一步提升指令跟随与对话能力,我们还探讨了基于理解模型训练聊天版本的方法,更契合人机实时交互需求。模型及聊天版已公开,包括模型权重、部分训练数据与推理代码,均发布于GitHub。

原文摘要 · Abstract (English)

With the development of Multimodal Large Language Models (MLLMs), numerous outstanding accomplishments have emerged within the open-source community. Due to the complexity of creating and training multimodal data pairs, it is still a computational and time-consuming process to build powerful MLLMs. In this work, we introduce Capybara-OMNI, an MLLM that trains in a lightweight and efficient manner and supports understanding text, image, video, and audio modalities. We present in detail the framework design, the data construction, and the training recipe, to develop an MLLM step-by-step to obtain competitive performance. We also provide exclusive benchmarks utilized in our experiments to show how to properly verify understanding capabilities across different modalities. Results show that by following our guidance, we can efficiently build an MLLM that achieves competitive performance among models of the same scale on various multimodal benchmarks. Additionally, to enhance the multimodal instruction following and conversational capabilities of the model, we further discuss how to train the chat version upon an MLLM understanding model, which is more in line with user habits for tasks like real-time interaction with humans. We publicly disclose the Capybara-OMNI model, along with its chat-based version. The disclosure includes both the model weights, a portion of the training data, and the inference codes, which are made available on GitHub.

多模态大模型开源语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。