通过熵引导分层策略,提升大模型推理强化学习的稳定性与效率。
Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

- 基于路径熵分层重放缓存,动态构建高对比度训练组。
- 在复杂推理任务上超越强基线,显著缓解熵坍缩问题。
- 适合追求稳定高效强化学习微调的研究者和工程师。
尽管强化学习能有效激励大语言模型的推理能力,但现有流程受限于训练不稳定和熵快速坍缩。这些问题通常源于标准采样中的“回滚静音”和低质量梯度信号。本文提出一种稳健的数据驱动框架以稳定强化学习训练。首先引入潜在感知查询挖掘(PAQM),动态筛选数据,聚焦于高潜力能力激发区——“提炼区”。其次提出混合分层回放(HSR),按路径熵(一种滚动置信度代理)和结果奖励对回滚样本进行分层重构。每个优化步骤中,HSR复用当前策略的“稳定锚点”和“困难负样本”,构建高对比度优化组,随后清空缓冲区进入下一步。该方法在有限算力下缓解熵坍缩,提升学习信号利用率。实验表明,本方法在复杂推理任务上优于多个强基线,为稳定高效的强化学习微调提供了原则性解决方案。
原文摘要 · Abstract (English)
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from "Rollout Silencing" and low-quality gradient signals in standard sampling procedures. In this work, we propose a robust, data-centric framework to stabilize RL training. We first introduce Potential-Aware Query Mining (PAQM), which filters data dynamically to focus on the "Distillation Zone"---samples with high potential for capability elicitation. Furthermore, we present Hybrid Stratified Replay (HSR), a novel mechanism that restructures batches by stratifying rollouts based on Path Entropy, a rollout-level confidence proxy, and outcome reward. Within each optimization step, HSR reuses current-policy "Stability Anchors" and "Hard Negatives" to construct high-contrast optimization groups, then clears its buffers before the next step. This approach mitigates entropy collapse while improving the utilization of learning signals under limited compute. Our method outperforms strong baselines on complex reasoning tasks, offering a principled solution for stable and efficient RL fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。