开源多模态专家混合模型Aria,性能媲美顶级闭源模型
Aria: An Open Multimodal Native Mixture-of-Experts Model
- 采用专家混合架构,视觉与文本每标记激活约3.9B和3.5B参数
- 在多模态、语言和编码任务上超越Pixtral-12B和Llama3.2-11B
- 提供完整代码与权重,支持实际应用中的快速部署与定制
信息以多样模态呈现。多模态原生AI模型对于整合现实世界信息并实现全面理解至关重要。尽管存在专有模型,但其封闭性阻碍了应用与适配。为此,我们提出Aria——一个开源的多模态原生模型,在广泛多模态、语言及编程任务中表现领先。Aria为专家混合模型,每视觉标记激活3.9B参数,每文本标记激活3.5B参数。其性能超越Pixtral-12B和Llama3.2-11B,且在多种多模态任务上可比肩顶尖闭源模型。我们通过四阶段预训练流程从头训练Aria,逐步提升语言理解、多模态理解、长上下文处理及指令遵循能力。现开放模型权重与代码库,便于在真实场景中便捷采用与适配。
原文摘要 · Abstract (English)
Information comes in diverse modalities. Multimodal native AI models are essential to integrate real-world information and deliver comprehensive understanding. While proprietary multimodal native models exist, their lack of openness imposes obstacles for adoptions, let alone adaptations. To fill this gap, we introduce Aria, an open multimodal native model with best-in-class performance across a wide range of multimodal, language, and coding tasks. Aria is a mixture-of-expert model with 3.9B and 3.5B activated parameters per visual token and text token, respectively. It outperforms Pixtral-12B and Llama3.2-11B, and is competitive against the best proprietary models on various multimodal tasks. We pre-train Aria from scratch following a 4-stage pipeline, which progressively equips the model with strong capabilities in language understanding, multimodal understanding, long context window, and instruction following. We open-source the model weights along with a codebase that facilitates easy adoptions and adaptations of Aria in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。