arXiv:2502.15803cs.LGcs.CL2025-02被引 6

轻量级多模态模型Megrez-3B-Omni,支持设备端图文音理解

Megrez-Omni Technical Report

  • 软硬协同设计,实现快速推理与小体积部署
  • 单模型在三类模态上均达当前最佳性能
  • 适合边缘设备部署,兼顾精度与效率

本文提出Megrez系列模型,包括语言模型Megrez-3B-Instruct和多模态模型Megrez-3B-Omni。通过软硬件协同设计,实现快速推理、模型紧凑与边缘智能的高效结合。Megrez-3B-Instruct具备高准确率、高速度和易用性,应用广泛。在此基础上构建的Megrez-3B-Omni是支持图像、文本、音频分析的设备端多模态大模型,在三类模态上均达到当前最优性能,展现出强泛化能力与鲁棒性,为多模态AI在边缘侧落地树立新标杆。

原文摘要 · Abstract (English)

In this work, we present the Megrez models, comprising a language model (Megrez-3B-Instruct) and a multimodal model (Megrez-3B-Omni). These models are designed to deliver fast inference, compactness, and robust edge-side intelligence through a software-hardware co-design approach. Megrez-3B-Instruct offers several advantages, including high accuracy, high speed, ease of use, and a wide range of applications. Building on Megrez-3B-Instruct, Megrez-3B-Omni is an on-device multimodal understanding LLM that supports image, text, and audio analysis. It achieves state-of-the-art accuracy across all three modalities and demonstrates strong versatility and robustness, setting a new benchmark for multimodal AI models.

多模态边缘计算轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。