通过多阶段微调提升多模态大模型性能,低成本实现更强通用能力。
MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- 基于信息密度与树状标签生成高质量跨域数据
- 多目标强化学习提升推理与响应简洁性
- 适合工业部署的高效训练与推理优化方案
我们提出MindGPT-4ov,一种通过多阶段后训练范式增强的多模态大语言模型(MLLM),涵盖数据生成、模型训练与高效部署。该方法在多个基准测试中达到最先进水平,且成本低廉,显著提升基础能力与泛化性能。核心创新包括:(1) 基于信息密度的数据生成方案,结合双维树状标签系统,实现高质量跨域数据的自动化生成;(2) 协作式课程监督微调,平衡领域知识注入与通用能力保留;(3) 混合强化学习框架,在提升推理能力的同时兼顾多样性探索、多模态感知保持与回答简洁性。此外,通过5D并行训练、算子优化与推理量化等基础设施改进,显著提升训练与推理效率,降低领域适配成本。实验表明,MindGPT-4ov在MMBench、MMStar、MathVision和MathVista等基准上优于现有模型,并在垂直领域任务中展现更优用户体验,支持从学术研究到工业部署的无缝迁移。该模型基于Qwen3-VL,相关权重、数据集与代码将近期开源,助力社区发展多模态大模型。
原文摘要 · Abstract (English)
We present MindGPT-4ov, a multimodal large language model (MLLM) that introduces a general post-training paradigm spanning data production, model training, and efficient deployment. It achieves state-of-the-art performance across multiple benchmarks at low cost, effectively enhancing the foundational capabilities of MLLMs and the generalization ability. Focusing on data construction, supervised fine-tuning strategies, and multimodal reinforcement learning methods, this work proposes three key innovations: (1) An information density-based data generation scheme, integrated with a dual-dimensional tree-structured label system, enabling automated generation of high-quality cross-domain data. (2) A collaborative curriculum supervised fine-tuning approach that balances the injection of domain-specific knowledge with the preservation of general capabilities. (3) A hybrid reinforcement learning paradigm that enhances reasoning ability while simultaneously addressing multi-objective optimization such as diversity exploration, maintenance of multimodal perception, and response conciseness. Moreover, we implement a series of infrastructure optimizations, such as 5D parallel training, operator optimization, and inference quantization to enhance training and inference efficiency while reducing the cost of domain adaptation. Experimental results demonstrate that the MindGPT-4ov model outperforms state-of-the-art models on benchmarks such as MMBench, MMStar, MathVision, and MathVista. In addition, MindGPT-4ov also demonstrates superior user experience in vertical domain tasks, enabling a seamless transition from academic research to industrial deployment. MindGPT-4ov provides a general post-training paradigm applicable to a wide range of MLLMs. The model weights, datasets, and code for the Qwen3-VL-based variants will be recently open-sourced to support the community's development of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。