arXiv:2512.00042cs.CVcs.AI2025-12

用高质量数据微调视觉语言模型,准确率逼近顶级闭源模型。

Closing the Gap: Data-Centric Fine-Tuning of Vision Language Models for the Standardized Exam Questions

  • 构建1.6亿词的多模态数据集,融合教材题解与课程图示
  • 在1854道标准化试题上达78.6%准确率,仅比Gemini低1.0%
  • 强调数据结构与表达语法对多模态推理的关键作用

多模态推理已成为现代AI研究的核心。标准化考试题目为这类推理提供了严谨的测试环境,具有结构化视觉上下文和可验证答案。尽管近期进展多集中于算法创新(如强化学习),但多模态推理的数据基础仍被忽视。本文表明,使用高质量数据进行监督微调(SFT)可媲美闭源方法。为此,我们构建了一个包含161.4百万词的多模态数据集,整合教科书题目-解答对、课程对齐图示及上下文材料,并采用优化的推理语法(QMSA)微调Qwen-2.5VL-32B模型。该模型在新发布的基准YKSUniform上达到78.6%准确率,该基准标准化了1,854道跨309个课程主题的多模态试题,仅比Gemini 2.0 Flash低1.0%。结果表明,数据构成与表征语法在多模态推理中起决定性作用。本工作建立了一种以数据为中心的框架,证明精心构建且课程对齐的多模态数据可使监督微调接近最先进水平。

原文摘要 · Abstract (English)

Multimodal reasoning has become a cornerstone of modern AI research. Standardized exam questions offer a uniquely rigorous testbed for such reasoning, providing structured visual contexts and verifiable answers. While recent progress has largely focused on algorithmic advances such as reinforcement learning (e.g., GRPO, DPO), the data centric foundations of vision language reasoning remain less explored. We show that supervised fine-tuning (SFT) with high-quality data can rival proprietary approaches. To this end, we compile a 161.4 million token multimodal dataset combining textbook question-solution pairs, curriculum aligned diagrams, and contextual materials, and fine-tune Qwen-2.5VL-32B using an optimized reasoning syntax (QMSA). The resulting model achieves 78.6% accuracy, only 1.0% below Gemini 2.0 Flash, on our newly released benchmark YKSUniform, which standardizes 1,854 multimodal exam questions across 309 curriculum topics. Our results reveal that data composition and representational syntax play a decisive role in multimodal reasoning. This work establishes a data centric framework for advancing open weight vision language models, demonstrating that carefully curated and curriculum-grounded multimodal data can elevate supervised fine-tuning to near state-of-the-art performance.

视觉语言模型多模态推理数据增强教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。