arXiv:2508.01540cs.CVcs.AI2025-08被引 1

轻量化视觉模型让手机运行多模态智能,功耗降41.1%

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning

  • 用少于1亿参数的轻量编码器+动态分辨率生成图像标记
  • 通过渐进式训练提升模型在多个子任务上的表现
  • 适合移动端部署,实测功耗降低41.1%,精度媲美顶尖模型

近年来,视觉语言模型(VLMs)取得了显著进展,广泛应用于日常生活。然而,其庞大的计算与存储需求给手机等移动设备的高效部署带来挑战。本文提出MagicVL-2B,专为旗舰智能手机优化的新型VLM。该模型采用参数少于100M的轻量视觉编码器,并设计动态分辨率机制,自适应生成图像标记,避免过度修改图像尺寸。为提升紧凑编码器在VLM中的性能,提出多模态课程学习策略,逐步增加任务难度与数据信息密度。大量实验表明,MagicVL-2B在标准VLM基准上达到当前顶尖模型的准确率,同时将本地设备功耗降低41.1%。结果证明,MagicVL-2B是实现真实移动场景下多模态智能的实用且可靠方案。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved remarkable breakthroughs in recent years, enabling a diverse array of applications in everyday life. However, the substantial computational and storage demands of VLMs pose significant challenges for their efficient deployment on mobile devices, which represent the most ubiquitous and accessible computing platforms today. In this work, we introduce MagicVL-2B, a novel VLM meticulously optimized for flagship smartphones. MagicVL-2B leverages a lightweight visual encoder with fewer than 100M parameters and features a redesigned dynamic resolution scheme that adaptively generates image tokens without excessive modification of image dimensions. To further enhance the performance of this compact encoder within VLMs, we propose a multimodal curriculum learning strategy that incrementally increases task difficulty and data information density throughout training. This approach substantially improves the model's performance across a variety of sub-tasks. Extensive evaluations on standard VLM benchmarks demonstrate that MagicVL-2B matches the accuracy of current state-of-the-art models while reducing on-device power consumption by 41.1%. These results establish MagicVL-2B as a practical and robust solution for real-world mobile vision-language applications, enabling advanced multimodal intelligence to run directly on smartphones.

视觉语言模型移动端部署轻量化课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。