调整多模态训练数据顺序,能显著影响模型理解、推理与文字识别能力的平衡。
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs

- 通过控制训练阶段的数据顺序,对比四种调度策略。
- 循序渐进训练使结构化推理最强,整体性能最优。
- 先学通用理解再引入文本密集任务,收敛更快更稳定。
近期多模态大模型在通用视觉理解、图表推理和文档感知方面表现优异,但其训练依赖异构监督数据,任务结构与学习需求差异大,而数据时间组织的影响仍不明确。本文通过固定骨干网络、可训练模块与优化流程,仅改变后对齐阶段数据的时序安排,系统研究数据组织对通用理解、结构化推理与细粒度OCR/文档理解之间权衡的影响。对比直接混合、课程学习、均衡采样与逆课程四种策略,在通用视觉指令遵循、图表推理、场景文本问答及文档问答任务上发现:数据组织是多模态适配的关键设计变量。课程学习在整体性能与结构化推理上表现最佳;均衡采样提升OCR能力但削弱综合平衡性;逆课程在最终性能与优化稳定性上最差。动态分析表明,先建立通用理解与推理能力,再引入密集文本任务,可实现更平滑优化与更快收敛。结果强调数据调度应作为多模态模型适应的显式设计维度。
原文摘要 · Abstract (English)
Recent multimodal large language models (MLLMs) perform strongly on general visual understanding, diagram and chart reasoning, and document-centric perception. However, these abilities are learned from heterogeneous supervision sources with very different task structures and learning demands, and the effect of their temporal organization during training remains underexplored. We study whether data organization affects the trade-off among general understanding, structured reasoning, and fine-grained OCR/document understanding in multimodal instruction tuning. To isolate this factor, we use a controlled three-stage training framework in which the backbone, trainable modules, and optimization pipeline are fixed across all runs, and only the temporal arrangement of post-alignment supervision is changed. We compare four strategies: direct mixture, curriculum training, balanced sampling, and reverse curriculum. Experiments on general visual instruction following, diagram reasoning, chart reasoning, scene-text question answering, and document question answering show that data organization is a first-order design variable in multimodal adaptation. Curriculum training gives the best overall trade-off and the strongest structured reasoning performance. Balanced sampling is better for OCR-oriented capability but weakens the broader capability balance. Reverse curriculum performs worst in both final performance and optimization stability. Training-dynamics analysis further suggests that building general understanding and reasoning before introducing OCR-intensive supervision leads to smoother optimization and faster convergence. These findings highlight data scheduling as an explicit design dimension for multimodal model adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。