发现扩散语言模型推理集中在动态混乱区,提升训练效率
Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones
- 识别出推理中的高不确定性动态区域,动态聚焦梯度更新
- 新方法在多个基准上提升推理准确率,且无需额外算力
- 适合研究扩散模型与强化学习融合的从业者参考
扩散大语言模型(dLLMs)正迅速发展为复杂推理的重要范式,强化学习被广泛用于下游对齐。现有基于轨迹的强化学习方法对去噪步骤均匀分配策略梯度,隐含假设所有步骤同等重要。我们通过熵不确定性、置信度-边际(CM)不确定性和熵变化率(RoEC)等步骤级指标分析轨迹,发现存在结构化的“混乱区”:短暂的不确定性高峰强烈预示最终成功或失败,而多数步骤保持稳定。为此提出自适应轨迹策略优化(ATPO),一种轻量级步骤选择策略,动态将梯度更新重分配至高杠杆步骤,不改变强化学习目标、奖励或计算预算。采用混合的RoEC+CM规则,ATPO在多个基准上显著提升推理准确率与训练稳定性,表明利用轨迹动态是推进dLLM强化学习的关键。
原文摘要 · Abstract (English)
Diffusion Large Language Models (dLLMs) are rapidly emerging alongside autoregressive models as a powerful paradigm for complex reasoning, with reinforcement learning increasingly used for downstream alignment. Existing trajectory-based RL methods uniformly allocate policy gradients across denoising steps, implicitly treating all steps as equally important. We challenge this assumption by analyzing trajectories with several step-level metrics: entropy-based uncertainty, Confidence-Margin (CM) uncertainty, and Rate of Entropy Change (RoEC). These reveal structured "zones of confusion": transient spikes in uncertainty and instability that strongly predict final success or failure, while most steps remain stable. We propose Adaptive Trajectory Policy Optimization (ATPO), a lightweight step-selection strategy that dynamically reallocates gradient updates to these high-leverage steps without changing the RL objective, rewards, or compute budget. Using a hybrid RoEC+CM rule, ATPO delivers substantial gains in reasoning accuracy and training stability across benchmarks, showing that exploiting trajectory dynamics is key to advancing dLLM RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。