arXiv:2603.04415cs.CLcs.CV2026-03

提出双调优框架,智能筛选适合推理训练的多模态数据

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training

  • 通过双重评估机制判断数据是否适合推理训练
  • 实验证明可识别有益、有害及适配直接作答的数据
  • 为多模态大模型提供精准训练策略匹配方案

推理后训练能提升大语言模型在数学与编程等复杂任务上的表现,但在多样化多模态任务中的效果仍不明确。当前主流团队推出的「指令」与「思维」并行模型既耗资源又难使用。已有研究指出,推理训练收益受基础模型能力、任务特性及链式思维(CoT)数据质量影响,但尚无系统性标准判断何时应采用推理训练及其适用数据。本文提出双调优(Dual Tuning)框架,针对目标任务与基础模型,联合评估训练数据是否有益,以及当前CoT内容能否使推理训练优于非推理方式。我们在空间、数学及跨学科任务上应用该方法,进一步分析强化学习与思维模式对推理效能的影响。结果可指导数据筛选:识别出利于推理训练的数据、更适合直接作答的数据,以及在两种训练模式下均不利的数据。本工作提供了选择合适训练数据与匹配后训练策略的量化依据。

原文摘要 · Abstract (English)

Reasoning post-training improves Large Language Models (LLMs) on complex tasks such as mathematics and coding, but its benefits across diverse multimodal tasks remains uncertain. The trend of releasing parallel "Instruct" and "Thinking" models by leading teams is both resource-intensive and user-unfriendly. Prior work finds that the gains from reasoning training are influenced by multiple factors, such as base model capabilities, task characteristics, and Chain-of-Thought (CoT) data quality. However, principled criteria for determining when reasoning post-training is beneficial and which data should support it are still lacking. In this paper, we propose Dual Tuning, a reasoning efficacy-driven data curation framework for multimodal LLMs training. Given a target task and a base model, Dual Tuning jointly evaluates whether the training data is beneficial and whether reasoning training with current CoT content yields positive gains over non-reasoning alternatives. We apply Dual Tuning across spatial, mathematical, and multi-disciplinary tasks, and further analyze how reinforcement learning and thinking patterns affect reasoning efficacy. The Dual Tuning results guide data curation by identifying data that benefit reasoning training, data better suited to direct-answer training, and data that are detrimental under both training modes. Our work provides quantitative criteria for selecting appropriate training data and matching post-training strategies.

多模态推理训练数据筛选双调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。