arXiv:2510.10201cs.LGcs.AI2025-10

用隐空间流动场构建奖励信号,提升大模型推理能力。

RLFR: Extending Reinforcement Learning for LLMs with Flow Environment

  • 从隐空间流动场提取速度偏差作为奖励信号
  • 在多模态与语言推理任务中表现可靠
  • 适合研究大模型奖励设计与隐空间挖掘的学者

基于可验证奖励的强化学习(RLVR)近期成为提升大语言模型推理能力的有前景框架。然而,仅依赖二元验证的策略容易忽略推理轨迹中的潜在价值探索。由于高质量过程奖励模型(PRM)标注成本高,现有工作尝试利用来自对数空间的熵和似然等辅助信号进行奖励塑造。本文提出一种新视角:基于隐空间的流动奖励(RLFR),通过离线高质量数据或在线拒绝采样数据构建模型隐空间的流动场,并量化策略隐状态在其中的速度偏差作为奖励信号。实验表明,良好的流动场可作为可靠的奖励收集环境,揭示了隐空间表达能力远未被充分挖掘。此外,RLFR能将任意离线专家数据压缩为参考,用于构建奖励信号,且其有效利用隐藏状态中的上下文依赖性,而非逐标记理解上下文。在语言与多模态推理基准上的实验验证了流动奖励的可靠性,提示了一种具有前景的辅助信号奖励塑造范式。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a promising framework for improving reasoning abilities in Large Language Models (LLMs). However, policy optimized with binary verification prone to overlook potential valuable exploration in reasoning trajectory. In view of heavy annotation cost of golden Process Reward Models (PRMs), recent works attempt using auxiliary signals for reward shaping of process tokens, involving entropy and likelihood collected from logit space. In this work, we offer a novel perspective on shaping RLVR with flow rewards derived from latent space, and propose RLFR, where the flow fields of model latents are constructed from either off-policy high-quality data and on-policy rejection sampling data, and the velocity deviations of policy latents within it are quantified to serve as a reward signal. RLFR first demonstrates that a well-established flow field can be a sound environment for reward signal collection, highlighting the expressive latent space is much underexplored. Moreover, RLFR is able to compress any off-policy expert data as reference for constituting reward signals, and we show that the efficient context dependence compressed within the hidden states are utilized, rather than individual token-level denotation for context comprehending. Experiments on both language and multimodal reasoning benchmarks demonstrate the reliability of flow rewards, and suggesting a promising paradigm for reward shaping with auxiliary signals.

强化学习大模型推理隐空间奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。