arXiv:2506.02308cs.LGcs.AI2025-06被引 3

按模态交互类型分组训练,提升多模态模型泛化能力

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

  • 根据模态间交互类型分组任务,如共享信息、独特信息或协同融合
  • 在多个基准上性能超越现有分组方法,平衡通用性与专属性
  • 适合需要高效多模态指令微调的科研与工业应用

近期多模态基础模型在多项任务中达到领先水平,主要得益于利用大规模无标签多模态数据的预训练范式,随后在精心筛选的标注数据集和高质量提示上进行指令微调。尽管指令微调正朝着更大规模的数据集扩展,但我们的研究发现,单纯增加指令微调任务数量并不能持续提升性能。相反,我们观察到:按模态间的共同交互方式(如发现冗余共享信息、优先选择具有独特信息的模态,或需协同融合以挖掘新信息)对任务进行分组,能促使模型在组内学习可迁移技能,同时抑制不匹配任务带来的干扰。为此,我们提出MINT——一种基于多模态交互类型的任务分组策略。实验表明,该方法显著优于现有的任务分组基线,在多模态指令微调中实现了通用性与专属性的高效平衡。

原文摘要 · Abstract (English)

Recent advances in multimodal foundation models have achieved state-of-the-art performance across a range of tasks. These breakthroughs are largely driven by new pre-training paradigms that leverage large-scale, unlabeled multimodal data, followed by instruction fine-tuning on curated labeled datasets and high-quality prompts. While there is growing interest in scaling instruction fine-tuning to ever-larger datasets in both quantity and scale, our findings reveal that simply increasing the number of instruction-tuning tasks does not consistently yield better performance. Instead, we observe that grouping tasks by the common interactions across modalities, such as discovering redundant shared information, prioritizing modality selection with unique information, or requiring synergistic fusion to discover new information from both modalities, encourages the models to learn transferrable skills within a group while suppressing interference from mismatched tasks. To this end, we introduce MINT, a simple yet surprisingly effective task-grouping strategy based on the type of multimodal interaction. We demonstrate that the proposed method greatly outperforms existing task grouping baselines for multimodal instruction tuning, striking an effective balance between generalization and specialization.

多模态指令微调任务分组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。