arXiv:2601.05870cs.LGcs.AI2026-01被引 2

用信息瓶颈机制让大模型推理更多样,避免陷入重复套路。

IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck

  • 通过信息瓶颈在高熵状态触发推理路径分叉,替代传统采样扰动。
  • 在4个数学推理基准上准确率提升最多5.3%,多样性指标提升7.4%。
  • 适合需要可靠、多样化推理的AI系统开发者或研究者。

基于可验证奖励的强化学习在大语言模型推理中的进展受到探索崩溃的持续阻碍:随机轨迹的语义同质性常使模型陷入狭窄且过度优化的行为。现有方法虽利用策略熵鼓励探索,但存在固有局限:全局熵正则易导致奖励滥用,引发无意义冗长;局部令牌选择性更新则难以克服预训练模型的强归纳偏见。为此,我们提出基于迭代信息瓶颈的潜在策略优化(IIB-LPO),将探索从词元分布的统计扰动转向推理轨迹的拓扑分支。IIB-LPO在高熵状态下触发潜在分支以拓展推理路径,并利用信息瓶颈原理作为轨迹过滤器与自奖励机制,确保探索的简洁性和信息量。在四个数学推理基准上的实证结果表明,IIB-LPO达到当前最优性能,准确率相比以往方法最高提升5.3%,多样性指标提升7.4%。

原文摘要 · Abstract (English)

Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Model (LLM) reasoning have been hindered by a persistent challenge: exploration collapse. The semantic homogeneity of random rollouts often traps models in narrow, over-optimized behaviors. While existing methods leverage policy entropy to encourage exploration, they face inherent limitations. Global entropy regularization is susceptible to reward hacking, which can induce meaningless verbosity, whereas local token-selective updates struggle with the strong inductive bias of pre-trained models. To address this, we propose Latent Policy Optimization via Iterative Information Bottleneck (IIB-LPO), a novel approach that shifts exploration from statistical perturbation of token distributions to topological branching of reasoning trajectories. IIB-LPO triggers latent branching at high-entropy states to diversify reasoning paths and employs the Information Bottleneck principle both as a trajectory filter and a self-reward mechanism, ensuring concise and informative exploration. Empirical results across four mathematical reasoning benchmarks demonstrate that IIB-LPO achieves state-of-the-art performance, surpassing prior methods by margins of up to 5.3% in accuracy and 7.4% in diversity metrics.

强化学习大模型推理探索多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。