arXiv:2509.18154cs.LGcs.CV2025-09被引 155

8B模型实现高效多模态推理,性能超越更大规模模型。

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

论文配图:MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
图 1 · 摘自论文原文
  • 采用3D-Resampler统一架构,压缩图像视频编码
  • 在视频和文档任务中表现领先,仅用46.7%显存
  • 适合需要低资源部署的多模态应用开发

多模态大语言模型(MLLMs)发展迅速,但训练与推理效率成为其普及的关键瓶颈。为此,我们提出参数量为8B的MiniCPM-V 4.5模型,通过三大改进提升效率与性能:统一的3D-Resampler架构实现图像与视频的高密度编码;无需复杂数据工程的统一学习范式,支持文档知识与文本识别;混合强化学习策略,兼顾短时与长时推理能力。OpenCompass评估显示,该模型在多项指标上超越GPT-4o-latest及更大的Qwen2.5-VL 72B等开源模型。尤其在视频理解任务中,于VideoMME基准上达到30B以下模型的顶尖水平,仅需46.7%显存占用和8.7%推理时间,显著优于Qwen2.5-VL 7B。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more accessible and scalable. To address the challenges, we present MiniCPM-V 4.5, an 8B parameter model designed for high efficiency and strong performance. We introduce three core improvements in model architecture, data strategy and training method: a unified 3D-Resampler model architecture for highly compact encoding over images and videos, a unified learning paradigm for document knowledge and text recognition without heavy data engineering, and a hybrid reinforcement learning strategy for proficiency in both short and long reasoning modes. Comprehensive experimental results in OpenCompass evaluation show that MiniCPM-V 4.5 surpasses widely used proprietary models such as GPT-4o-latest, and significantly larger open-source models such as Qwen2.5-VL 72B. Notably, the strong performance is achieved with remarkable efficiency. For example, on the widely adopted VideoMME benchmark, MiniCPM-V 4.5 achieves state-of-the-art performance among models under 30B size, using just 46.7\% GPU memory cost and 8.7\% inference time of Qwen2.5-VL 7B.

多模态高效推理视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。