筛选与模型匹配的难例数据,能显著提升多模态推理效果。
Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols
- 基于对齐数据集进行难度过滤,提升推理准确率。
- 1000条对齐数据内,难度筛选使性能优于随机采样。
- 该方法适合数据受限的多模态模型微调场景。
我们研究在固定微调协议下多模态推理的数据筛选策略,其中基础模型、优化器、训练调度和评估流程保持不变,唯一变量是训练数据。以NeurIPS 2025 DCVLR挑战赛为受控测试平台,分析源数据对齐性、模型相对难度、数据集规模、多样性启发式及重写合成混合数据对下游推理准确率的影响。在多种干预中,对齐语料库上的难度过滤在相同规模下带来最大提升。这一效果不仅由LiveXivTQA加权解释:按基准分解显示,大部分增益来自最大的非LiveXivTQA基准OlympiadBench。Qwen衍生的难度评分也部分迁移至其他模型族,但收益具架构依赖性。相反,超过约1000条对齐样本后,增加数据量主要降低运行间方差,而测试的多样性与重写CoSyn混合数据未超越难度过滤基线。这些结果为数据受限的多模态推理微调提供具体实证建议,而非跨所有训练范式的普适结论。
原文摘要 · Abstract (English)
We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule, and evaluation pipeline are held constant and the main degree of freedom is the training data. Using the NeurIPS 2025 Data Curation for Vision--Language Reasoning (DCVLR) challenge as a controlled testbed, we analyze how source-dataset alignment, model-relative difficulty, dataset size, diversity heuristics, and rewritten synthetic mixtures affect downstream reasoning accuracy. Among the tested interventions, difficulty filtering on an aligned source corpus provides the strongest gains at matched scale. The effect is not explained only by LiveXivTQA weighting: a per-benchmark decomposition shows that much of the improvement over random sampling comes from OlympiadBench, the largest non-LiveXivTQA benchmark. Qwen-derived difficulty scores also transfer to some additional model families, though the benefit is architecture-dependent. In contrast, increasing dataset size beyond roughly 1k aligned examples mainly reduces run-to-run variance under the fixed recipe, while the diversity and rewritten CoSyn mixtures we tested do not improve over the difficulty-filtered baseline. These results provide a scoped empirical recipe for data-constrained multimodal reasoning fine-tuning, rather than a universal claim about data selection across all training regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。