用自生成数据实现多模态模型自我进化,打破预训练瓶颈
Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition
- 通过自生成数据进行多模态预训练,构建可自我提升的认知系统
- 仅用21.3万自生成样本,在多个评测上提升3.5%以上
- 适合追求高效模型迭代与多模态推理能力的研究者
近年来,(多模态)大语言模型((M)LLMs)的发展重心从预训练转向推理时计算和后训练优化,主要由于高质量人工数据的稀缺。然而,仅靠这些策略难以带来显著性能提升。本文认为,模型进步需要预训练、推理时计算与后训练优化的深度融合。为此,提出自提升认知(SIcog)框架,通过自生成数据进行多模态预训练,赋予模型多模态知识与系统性认知能力。具体包括:分步视觉理解的描述链(Chain-of-Description),以及支持深度多模态推理的结构化思维链(CoT)。SIcog首先以极少量外部监督增强基础模型的感知与推理能力;随后,模型为无标签图像及图文对生成候选描述与CoT推理结果,并通过语义相似性引导的自一致性机制筛选高质量样本。这些高质自生成数据用于大规模多模态预训练,形成自我改进闭环。实验表明,仅使用21.3万自生成样本,SIcog在MMStar上提升+3.6%,AI2D上提升+3.5%,优于以往预训练方法;结合后训练的CoT技术,在MMVet上提升+9%,ScienceQA上提升+8.5%。
原文摘要 · Abstract (English)
Recent progress in (multimodal) large language models ((M)LLMs) has shifted focus from pre-training to inference-time computation and post-training optimization, largely due to concerns over the availability of high-quality human data. However, these strategies alone are insufficient to drive substantial model improvements. We argue that effective model advancement requires strong synergy among pre-training, inference-time computation, and post-training optimization. In this paper, we introduce Self-Improving cognition (SIcog), a self-learning framework for constructing next-generation foundation MLLMs by imparting multimodal knowledge and enhancing systematic cognitive capabilities through multimodal pre-training with self-generated data. Specifically, we propose Chain-of-Description for step-by-step visual understanding and integrate structured Chain-of-Thought (CoT) reasoning to support in-depth multimodal reasoning. SIcog first equips a base model with systematic perception and reasoning using minimal external supervision. The enhanced models then generate candidate image captions and CoT reasoning responses for unlabeled images and image-question pairs across diverse tasks, which are filtered through a semantic-similarity-guided self-consistency mechanism. These high-quality, self-generated samples enable large-scale multimodal pre-training, creating a self-improvement loop. Experiments demonstrate SIcog's effectiveness in developing MLLMs with enhanced multimodal cognition. Using only 213K self-generated pre-training samples, SIcog achieves significant improvements, including +3.6% on MMStar and +3.5% on AI2D, outperforming previous pre-training approaches. When combined with post-training techniques for CoT reasoning, SIcog yields +9% gains on MMVet and +8.5% on ScienceQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。