arXiv:2604.17892cs.LGcs.AI2026-04ACL被引 3

让大模型在连续空间中探索更多推理路径,提升强化学习效果

LEPO: Latent Reasoning Policy Optimization for Large Language Models

  • 用Gumbel-Softmax注入可控随机性,保持推理多样性
  • 直接对连续隐变量应用强化学习,性能超越现有方法
  • 适合需要多路径探索的复杂推理任务

最近,将连续空间中的潜在推理引入大语言模型(LLMs),以利用其中丰富的信息。然而,缺乏随机采样使得这些方法不可避免地退化为确定性推理,无法发现多样化的推理路径。为此,我们通过Gumbel-Softmax向潜在推理注入可控随机性,恢复了LLMs的探索能力,并增强了其与强化学习(RL)的兼容性。在此基础上,我们提出一种新框架——潜在推理策略优化(LEPO),直接对连续隐表示应用强化学习。具体而言,在采样阶段,LEPO保持随机性以实现多样化轨迹采样;在优化阶段,构建统一梯度估计,同时优化隐表示与离散词元。大量实验表明,LEPO显著优于现有的离散与潜在推理强化学习方法。

原文摘要 · Abstract (English)

Recently, latent reasoning has been introduced into large language models (LLMs) to leverage rich information within a continuous space. However, without stochastic sampling, these methods inevitably collapse to deterministic inference, failing to discover diverse reasoning paths. To bridge the gap, we inject controllable stochasticity into latent reasoning via Gumbel-Softmax, restoring LLMs' exploratory capacity and enhancing their compatibility with Reinforcement Learning (RL). Building on this, we propose \textbf{\underline{L}}atent R\textbf{\underline{e}}asoning \textbf{\underline{P}}olicy \textbf{\underline{O}}ptimization~(\textbf{LEPO}), a novel framework that applies RL directly to continuous latent representations. Specifically, in rollout stage, LEPO maintains stochasticity to enable diverse trajectory sampling, while in optimization stage, LEPO constructs a unified gradient estimation for both latent representations and discrete tokens. Extensive experiments show that LEPO significantly outperforms existing RL methods for discrete and latent reasoning.

强化学习大模型推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。