arXiv:2602.16742cs.LGcs.AI2026-02被引 3

构建10.3万条多模态数学数据,提升模型视觉推理能力

DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning

  • 基于可验证奖励的强化学习,构建覆盖广泛数学知识点的数据集
  • 模型在多模态数学基准上表现优异,且能泛化到通用多模态任务
  • 数据涵盖丰富视觉元素,适合训练具备强视觉推理能力的模型

基于可验证奖励的强化学习(RLVR)已被证明能有效提升大型多模态模型(LMMs)的视觉反思与推理能力。然而,现有数据集主要依赖小规模人工构造或已有资源的重组,导致数据多样性与覆盖面不足,制约了模型性能的进一步提升。为此,我们提出 extbf{DeepVision-103K},一个全面用于RLVR训练的多模态数学数据集,覆盖K12阶段多样数学主题、广泛知识要点及丰富的视觉内容。在该数据集上训练的模型在多模态数学基准上表现强劲,并能有效泛化至通用多模态推理任务。进一步分析表明,模型的视觉感知、反思与推理能力均得到显著增强,验证了DeepVision对推动多模态推理的有效性。数据可在 Hugging Face 获取。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has been shown effective in enhancing the visual reflection and reasoning capabilities of Large Multimodal Models (LMMs). However, existing datasets are predominantly derived from either small-scale manual construction or recombination of prior resources, which limits data diversity and coverage, thereby constraining further gains in model performance. To this end, we introduce \textbf{DeepVision-103K}, a comprehensive dataset for RLVR training that covers diverse K12 mathematical topics, extensive knowledge points, and rich visual elements. Models trained on DeepVision achieve strong performance on multimodal mathematical benchmarks, and generalize effectively to general multimodal reasoning tasks. Further analysis reveals enhanced visual perception, reflection and reasoning capabilities in trained models, validating DeepVision's effectiveness for advancing multimodal reasoning. Data: \href{https://huggingface.co/datasets/skylenage/DeepVision-103K}{this url}.

多模态推理数学理解强化学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。