arXiv:2505.13261cs.CV2025-05被引 9

通过难度先验提升强化学习微调的多模态推理能力

Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning

  • 用两阶段训练过滤简单和极难样本,提升梯度质量
  • 基于群体准确率动态重加权优势值,增强难题学习信号
  • 在复杂样本中加入难度提示,引导模型深度思考与验证

本文研究如何显式建模问题难度先验以提升基于强化学习的多模态推理微调效果。首先,通过离线数据筛选,利用基座模型多轮采样分析两个数据集的U型难度分布,并剔除过于简单或极端困难的提示,以提供有意义的梯度,随后进行两阶段训练。其次,实现在线优势差异化,以群体经验准确率为难度代理,自适应重加权优势估计,强化对复杂问题的学习信号。最后,在第二阶段对复杂样本引入难度提示,引导模型调整推理深度并执行反思性验证。该综合方法仅使用2K+0.6K两阶段训练数据,在多个多模态数学推理基准上表现显著。

原文摘要 · Abstract (English)

In this work, we investigate how explicitly modeling problem's difficulty prior information shapes the effectiveness of reinforcement learning based fine-tuning for multimodal reasoning. Our exploration mainly comprises of following three perspective: First, through offline data curation, we analyze the U-shaped difficulty distribution of two given datasets using the base model by multi-round sampling, and then filter out prompts that are either too simple or extremely difficult to provide meaningful gradients and perform subsequent two-stage training. Second, we implement an online advantage differentiation, computing group-wise empirical accuracy as a difficulty proxy to adaptively reweight advantages estimation, providing stronger learning signals for more challenging problems. Finally, we introduce difficulty hints as explicit prompts for more complex samples in the second training stage, encouraging the model to calibrate its reasoning depth and perform reflective validation checks. Our comprehensive approach demonstrates significant performances across various multi-modal mathematical reasoning benchmarks with only 2K+0.6K two-stage training data.

强化学习多模态推理难度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。