通过分拆困惑度空间,实现大模型推理中探索与利用的精细平衡。
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off

- 按困惑度将样本分为高探索、低探索区域,实现细粒度控制。
- 在不干扰验证奖励的前提下,双向分配奖励提升优化稳定性。
- 在数学推理和函数调用任务上显著提升大模型性能。
基于可验证奖励的强化学习(RLVR)推动了大语言模型(LLMs)推理能力的显著进步。然而,如何有效管理探索与利用之间的权衡仍是关键挑战。本文深入分析训练过程中极难与极易样本带来的探索-利用困境,提出一种新的细粒度权衡机制。具体而言,引入困惑度空间解耦策略,将样本空间划分为高困惑度(探索)与低困惑度(利用)子空间,从而挖掘需要精细权衡的样本。随后,设计一种双向奖励分配机制,在最小影响验证奖励的前提下,实现困惑度引导的探索与利用,促进更稳定的策略优化。最后,在主流的数学推理与函数调用任务上评估该方法,实验结果表明其显著优于现有方法,验证了通过细粒度探索-利用权衡提升大模型性能的有效性。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs). However, effectively managing the exploration and exploitation trade-off remains a critical challenge. In this paper, we fully analyze the exploration and exploitation dilemma of extremely hard and easy samples during the training and propose a new fine-grained trade-off mechanism. Concretely, we introduce a perplexity space disentangling strategy that divides the sample space into distinct exploration (high perplexity) and exploitation (low perplexity) subspaces, thereby mining fine-grained samples requiring exploration-exploitation trade-off. Subsequently, we propose a bidirectional reward allocation mechanism with a minimum impact on verification rewards to implement perplexity-guided exploration and exploitation, enabling more stable policy optimization. Finally, we have evaluated our method on two mainstream tasks: mathematical reasoning and function calling, and experimental results demonstrate the superiority of the proposed method, confirming its effectiveness in enhancing LLM performance by fine-grained exploration-exploitation trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。