开源多模态模型Moxin-VLM/VLA,支持视觉、语言与动作理解。
Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- 基于Moxin框架,完全开源训练数据与代码
- 在多任务评测中表现优于现有模型
- 适合研究者快速部署视觉-语言-动作系统
近年来,大型语言模型(LLMs)迅速发展,以GPT-4和GPT-o1为代表的专有模型广受关注,而开源模型如LLaMA和Mistral也推动了模型的广泛应用。本文提出完全开源的Moxin 7B模型,遵循模型开放性框架,实现训练过程、数据集与实现细节的全面透明,促进更开放协作的研究生态。为拓展能力,我们开发了三个变体:Moxin-VLM(视觉-语言)、Moxin-VLA(视觉-语言-动作)和Moxin-Chinese(中文能力)。实验表明,这些模型在多项评估中表现优异。训练采用开源框架与公开数据集,并已发布模型、数据及代码,支持复现与二次开发。
原文摘要 · Abstract (English)
Recently, Large Language Models (LLMs) have undergone a significant transformation, marked by a rapid rise in both their popularity and capabilities. Leading this evolution are proprietary LLMs like GPT-4 and GPT-o1, which have captured widespread attention in the AI community due to their remarkable performance and versatility. Simultaneously, open-source LLMs, such as LLaMA and Mistral, have made great contributions to the ever-increasing popularity of LLMs due to the ease to customize and deploy the models across diverse applications. Moxin 7B is introduced as a fully open-source LLM developed in accordance with the Model Openness Framework, which moves beyond the simple sharing of model weights to embrace complete transparency in training, datasets, and implementation detail, thus fostering a more inclusive and collaborative research environment that can sustain a healthy open-source ecosystem. To further equip Moxin with various capabilities in different tasks, we develop three variants based on Moxin, including Moxin-VLM, Moxin-VLA, and Moxin-Chinese, which target the vision-language, vision-language-action, and Chinese capabilities, respectively. Experiments show that our models achieve superior performance in various evaluations. We adopt open-source framework and open data for the training. We release our models, along with the available data and code to derive these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。