MM1.5通过数据优化提升多模态模型在图文理解等任务的表现。
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
- 基于数据驱动方法,系统优化预训练与微调阶段的数据组合。
- 1B~30B参数模型在小规模下仍达优秀性能,视频和UI特化版本表现突出。
- 适合关注多模态模型训练策略与数据设计的研究者参考。
我们提出MM1.5,一个新型多模态大语言模型系列,旨在增强文本丰富的图像理解、视觉指代与定位以及多图推理能力。基于MM1架构,MM1.5采用以数据为中心的训练方法,系统探索全训练周期中不同数据混合的影响:包括高质量OCR数据与合成描述用于持续预训练,以及优化后的视觉指令微调数据集用于监督微调。模型涵盖1B至30B参数,包含密集与混合专家(MoE)两种结构,实验证明即使在小规模(1B和3B)下,通过精心的数据筛选与训练策略也能实现优异性能。此外,我们推出了两个专用变体:MM1.5-Video用于视频理解,MM1.5-UI专为移动端界面理解设计。通过广泛的实证研究与消融实验,我们深入剖析了训练过程中的关键决策,为未来多模态大模型研发提供宝贵指导。
原文摘要 · Abstract (English)
We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。