arXiv:2409.05840cs.CL2024-09被引 47

用迭代演化提升多模态数据质量,让大模型更懂图与文。

MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

  • 通过感知、推理、交互三阶段迭代优化指令数据
  • 在13项任务上平均准确率提升3.1个百分点
  • 用更少数据达到9项任务的顶尖水平,适合模型训练者

多模态大语言模型(MLLMs)的发展在多模态代理、具身智能等领域需求推动下取得显著进展。尽管模型驱动方法尝试通过多样化架构提升能力,但收益逐渐边际化。相比之下,数据驱动方法虽更有效,却受限于数据多样性与复杂性。高质量数据缺失成为制约发展的重要瓶颈。为此,我们提出MMEvol——一种新型多模态指令数据演化框架。该框架通过精细感知、认知推理与交互演化相结合的方式,迭代提升数据质量,生成更具复杂性与多样性的图文指令数据集,从而增强MLLMs能力。基于初始指令集SEED-163K,MMEvol系统性扩展指令类型、延长视觉推理步骤以强化认知能力,并深入挖掘图像中的细粒度信息以提升视觉理解与鲁棒性。为全面评估效果,我们在13个视觉-语言任务上开展定性分析与定量实验。结果表明,相较于使用初始种子数据训练的基线模型,本方法平均准确率提升3.1个百分点;且在九项任务中以更少数据达到当前最优(SOTA)表现。

原文摘要 · Abstract (English)

The development of Multimodal Large Language Models (MLLMs) has seen significant advancements with increasing demands in various fields (e.g., multimodal agents, embodied intelligence). While model-driven approaches attempt to enhance MLLMs capabilities through diverse architectures, the gains have become increasingly marginal. Conversely, data-driven methods, which scale up image-text instruction data, are more effective but face limited data diversity and complexity challenges. The absence of high-quality data constitutes a significant development barrier for MLLMs. To address the data quality bottleneck, we propose MMEvol, a novel multimodal instruction data evolution framework. This framework iteratively improve data quality through a refined combination of fine-grained perception, cognitive reasoning, and interaction evolution, generating a more complex and diverse image-text instruction dataset that empowers MLLMs with enhanced capabilities. Beginning with an initial set of instructions, SEED-163K, we utilize MMEvol to systematically broaden the diversity of instruction types, extend visual reasoning steps to improve cognitive reasoning abilities, and thoroughly explore fine-grained information within images to enhance visual understanding and robustness. To comprehensively evaluate the effectiveness of our approach, we conduct extensive qualitative analysis and quantitative experiments across 13 vision-language tasks. Compared to baseline models trained with the initial seed data, the results demonstrate that our method achieves an average accuracy improvement of 3.1 percentage points. Furthermore, our approach reaches state-of-the-art (SOTA) performance in nine tasks using significantly less data compared to state-of-the-art models.

多模态数据演化大模型指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。