arXiv:2608.02867cs.CLcs.AI2026-08被引 1

探究大模型推理中真实思维分支是否被强化,发现强化学习反而抑制了多样性。

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

论文配图:BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
图 1 · 摘自论文原文
  • 通过构建语义等价的推理树结构,区分风格差异与真实推理分支。
  • RLVR模型在测试时的语义分支熵显著降低,但更遵守环境约束并提升回溯能力。
  • 揭示了强化学习提升采样效率的同时牺牲了思维多样性,适合研究推理机制者阅读。

尽管基于可验证奖励的强化学习(RLVR)在多种推理任务中提升了大语言模型(LLMs)的表现,但其究竟是拓展了推理能力边界,还是仅提高了采样效率,仍存在争议。本文通过受控迷宫求解实验,并基于数学推理轨迹提取语义等价的树结构(BODHI-Trees),研究了经RLVR训练的LLMs在测试时的探索性质。该方法帮助我们区分由风格差异引起的熵与真实的推断分支熵。研究发现,RLVR模型中观察到的策略熵坍塌并非仅是语法层面的,还伴随着语义分支熵的显著下降。虽然RLVR增强了对环境约束的遵守和回溯能力,却限制了后续路径的空间;我们提供的证据表明,这可能正是RLVR实现样本效率提升的原因,代价是真实推理路径多样性的减少。

原文摘要 · Abstract (English)

Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary, or just improves sampling efficiency. In this paper, we investigate the nature of test-time exploration in RLVR-trained LLMs by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence. This helps us delineate between entropy arising from stylistic variations and genuine inferential branching. Our findings demonstrate that the policy entropy collapse observed in RLVR models is not merely syntactic, and is accompanied by a significant reduction in semantic branching entropy. While RLVR improves adherence to environmental constraints and backtracking capabilities, it constricts the space of continuations; we provide evidence suggesting that this might be responsible for the sample efficiency gains of RLVR, albeit at the cost of genuine rollout diversity.

大模型推理强化学习思维链语义分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。