剖析大模型在强化学习中的探索机制,揭示其能力边界与优化路径。
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
- 构建量化指标,分析大模型探索空间的边界特性
- 发现熵与性能在训练各阶段存在动态权衡关系
- 提出将探索优势转化为实际性能提升的有效方法
基于可验证奖励的强化学习(RLVR)已成为提升大语言模型(LLM)推理能力的重要范式。与传统强化学习不同,RLVR利用规则化反馈引导LLM生成并优化复杂推理链——这一过程高度依赖有效的探索策略。尽管已有研究证明了RLVR的实证效果,但大模型探索行为的基本机制仍缺乏深入探讨。本技术报告系统性地研究了RLVR中的探索能力,涵盖四个方面:(1)探索空间塑造,开发量化指标刻画LLM的能力边界;(2)熵-性能权衡,分析训练阶段、个体实例及词元级模式下的变化规律;(3)强化学习性能优化,考察如何有效将探索收益转化为可测量的性能提升。通过整合已有发现与新实证证据,本文旨在为推进RLVR系统提供基础框架。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs). Unlike traditional RL approaches, RLVR leverages rule-based feedback to guide LLMs in generating and refining complex reasoning chains -- a process critically dependent on effective exploration strategies. While prior work has demonstrated RLVR's empirical success, the fundamental mechanisms governing LLMs' exploration behaviors remain underexplored. This technical report presents a systematic investigation of exploration capacities in RLVR, covering four main aspects: (1) exploration space shaping, where we develop quantitative metrics to characterize LLMs' capability boundaries; (2) entropy-performance exchange, analyzed across training stages, individual instances, and token-level patterns; and (3) RL performance optimization, examining methods to effectively translate exploration gains into measurable improvements. By unifying previously identified insights with new empirical evidence, this work aims to provide a foundational framework for advancing RLVR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。