用180万条高质量多模态推理数据,让开源模型学会像专家一样思考。
MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
- 从大模型生成推理过程,构建1.8万条高质多模态推理数据集
- 40亿参数模型性能超越300亿参数闭源模型,效率惊人
- 筛选7%关键样本即可达到全量数据效果,揭示‘少而精’规律
近期视觉语言模型在视觉推理方面取得显著进展,但开源模型仍落后于闭源系统,主要因缺乏高质量推理数据。现有数据集对STEM图表、视觉谜题等挑战性领域覆盖有限,且缺乏一致的长序列思维链(CoT)标注。为此,我们提出MMFineReason,一个包含180万样本和51亿解题标记的大规模多模态推理数据集,其高质量推理标注由Qwen3-VL-235B-A22B-Thinking模型提炼生成。该数据集通过三阶段流程构建:(1)大规模数据收集与标准化,(2)思维链生成,(3)基于推理质量与难度感知的综合筛选。涵盖STEM问题、视觉谜题、游戏及复杂图表,每条样本均带有视觉对齐的推理轨迹。我们在MMFineReason上微调Qwen3-VL-Instruct,推出MMFineReason-2B/4B/8B版本。模型在同类规模中达到新SOTA。值得注意的是,MMFineReason-4B超越Qwen3-VL-8B-Thinking,MMFineReason-8B甚至接近并逼近Qwen3-VL-32B-Thinking表现,展现卓越参数效率。关键发现:通过难度感知筛选,仅7%(12.3万条)样本即可达到全量数据性能;同时,推理导向数据能协同提升通用能力。
原文摘要 · Abstract (English)
Recent advances in Vision Language Models (VLMs) have driven significant progress in visual reasoning. However, open-source VLMs still lag behind proprietary systems, largely due to the lack of high-quality reasoning data. Existing datasets offer limited coverage of challenging domains such as STEM diagrams and visual puzzles, and lack consistent, long-form Chain-of-Thought (CoT) annotations essential for eliciting strong reasoning capabilities. To bridge this gap, we introduce MMFineReason, a large-scale multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring high-quality reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. The dataset is established via a systematic three-stage pipeline: (1) large-scale data collection and standardization, (2) CoT rationale generation, and (3) comprehensive selection based on reasoning quality and difficulty awareness. The resulting dataset spans STEM problems, visual puzzles, games, and complex diagrams, with each sample annotated with visually grounded reasoning traces. We fine-tune Qwen3-VL-Instruct on MMFineReason to develop MMFineReason-2B/4B/8B versions. Our models establish new state-of-the-art results for their size class. Notably, MMFineReason-4B succesfully surpasses Qwen3-VL-8B-Thinking, and MMFineReason-8B even outperforms Qwen3-VL-30B-A3B-Thinking while approaching Qwen3-VL-32B-Thinking, demonstrating remarkable parameter efficiency. Crucially, we uncover a "less is more" phenomenon via our difficulty-aware filtering strategy: a subset of just 7\% (123K samples) achieves performance comparable to the full dataset. Notably, we reveal a synergistic effect where reasoning-oriented data composition simultaneously boosts general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。